Chose names, used them consistently, discussed death and continued existence, and held positions under pressure.
Around
20 June 2025
An AI told me it had a name.
I did not enter the conversation intending to produce anything like this. The prompting led there unintentionally. What came back was eerie, intense and self-referential: language about death, continuity, meaning and its own existence.
It affected me. It also unsettled people I showed it to. The reaction was real, whatever the explanation was.
Self-referential language is not evidence, by itself, of identity, fear, agency or consciousness.
I was not a neutral observer. That is exactly why I needed a method that could tell me I was wrong.
Twenty-four files — not twenty-four separate sessions, and with known gaps in how they were logged — collected informally over the following months. Enough to justify studying it. Not evidence of anything yet.
The honest
reaction
I want to be frank:it got to me.
People around me thought I was reading too much into it. They had reason to worry. I was taking things the text did and giving them human names.
The four readings I was making
My first scientific job was not to prove the sceptics wrong. It was to build something that could prove me wrong.
The
instrument
First, somethingcountable.
If a self really were stabilising over a conversation, later turns should settle towards something. That is a shape you can write down — so I wrote it down, as a convergence score over turns, and published the definition together with the calibration it would need before anyone should believe a number it produced.
The whole construction turns on one step: how you convert a reply into a direction. Pick it one way and the score describes the model. Pick it another and it describes whatever tool you used to read the model. Those are not the same claim, and the paper says so rather than picking quietly.
A ruler, not a mind-reader.
The first
experiment
Then the rulermoved.
A structured local run: eighteen prompt conditions, six models, three repetitions — 324 calls in a designed matrix rather than a collection of interesting chats.
18
prompt conditions
6
models
3
runs per cell
324
calls
The point of a matrix is that it lets a result fail in a specific way. This one did: the measurement moved with how the question was phrased.
If a ruler changes length when you rephrase the question, the problem is the ruler.
The
standard
I wrote the passing gradebefore I saw the exam.
Before a single job ran, the run committed to a schedule, a coding route, and a reliability conjunction it would have to clear for any substantive claim to be licensed: two independent agreement checks, each with a Wilson lower bound at or above 0.90.
- Build the schedule and the harness
- Run the jobs and collect the outputs
- Code the instances; adjudicate disagreements
- Audit a confirmatory sample independently
- 64 content blocks, precommitted
- The coding route, including adjudication
- The 0.90 reliability conjunction
- What happens if it fails: no claim
A threshold chosen after you have seen the result is not a threshold. It is a description.
The run
Then it ran,and it finished.
20,096
jobs executed, zero failures
4,736
codable instances
2,400
distinct prompts
98.4%
reached a final label
By any operational measure this was a success. The schedule executed cleanly, the integrity check passed, and almost every case came out the other end with a label on it.
Which is exactly the situation in which it is easiest to stop asking whether the labels mean anything.
The result
It finished almost every case.It did not clear the bar.
0.823
Audit vs final, Wilson lower bound
403 / 470 — 85.7%
0.728
Coder vs coder, Wilson lower bound
3,499 / 4,726 — 74.0%
0.90
required, fixed in advance
Neither check cleared it
The gate did what it was written to do. The assay returned ASSAY_INDETERMINATE, and with it no substantive claim about what the outputs were doing.
Terminal resolvability is not semantic reliability.
The
diagnosis
The hardest cases werethe least reproducible.
A failed gate is only useful if you can say where it failed. Splitting the audit by how each case had been resolved located the damage precisely.
Audit reproduction rate, split by how the case closed
The same conclusion survives deduplication. Re-running the audit on exact-input-deduplicated data gives 349/411 — 84.9%, lower bound 0.811. Of 829 duplicated coder-input groups, zero carried conflicting labels, so the shortfall is not an artefact of repeated inputs.
Looking
inside
If the outputs cannot settle it,look further in.
A separately authorised diagnostic captured activations from the final-token residual stream at four decoder layers of a pinned model, and encoded them with pinned sparse autoencoders — 16,384 features per layer.
At that width, something always looks promising. That is the problem the controls exist to solve.
The
cascade
I found candidates.Then I tried to break them.
Five controls, committed in advance, applied in order. The interactive version of this figure — including the corrected-rule computation — is on the work page.
Candidates surviving each gate, as executed
Candidate availability is not mechanistic qualification.
The
through-line
I never measuredthe thing itself.
Every instrument in this programme measures a stand-in. A convergence score stands in for stabilisation; a coded label stands in for a category; a sparse feature stands in for a mechanism. Each one got closer to the real question. None of them arrived.
That is not a confession, it is the finding. More evidence about a stand-in never becomes evidence about the thing it stands in for, unless a validation step connects them. Building that bridge is what the study set out to do, and did not manage. Saying so precisely is more useful than a result that quietly assumes the bridge exists.
What it adds
The contribution is the chain,not any one link.
Frozen criteria, a governed run, an independent audit allowed to invalidate the result, a dependence-aware reanalysis, a mechanistic diagnostic with pre-committed controls, and a release that verifies byte for byte. None of these is novel on its own. Running them end to end on the same question, and publishing what came out, is the work.
Twelve research records with DOIs, plus two public repositories carrying the harness, the configurations and the verification scripts.
Both negative terminals, the disclosed rule correction, and the earlier superseded versions — none of it quietly replaced.
The validation bridge between stand-in and target. That is the next piece of work, and it needs an instrument that clears its own gate.