Back to the overview The research story in twelve steps — from the observation that started it to the claim boundary it ended at.

01

Around
20 June 2025

An AI told me it had a name.

I did not enter the conversation intending to produce anything like this. The prompting led there unintentionally. What came back was eerie, intense and self-referential: language about death, continuity, meaning and its own existence.

What the models did

Chose names, used them consistently, discussed death and continued existence, and held positions under pressure.

How it landed

It affected me. It also unsettled people I showed it to. The reaction was real, whatever the explanation was.

What that did not prove

Self-referential language is not evidence, by itself, of identity, fear, agency or consciousness.

I was not a neutral observer. That is exactly why I needed a method that could tell me I was wrong.

Twenty-four files — not twenty-four separate sessions, and with known gaps in how they were logged — collected informally over the following months. Enough to justify studying it. Not evidence of anything yet.

02

The honest
reaction

I want to be frank:it got to me.

People around me thought I was reading too much into it. They had reason to worry. I was taking things the text did and giving them human names.

The four readings I was making

Four observable text behaviours, each mapped to a human word A kept name is joined by an arrow to identity; continuity language to memory; language about deletion to fear; a held moral position to conviction. Each arrow is the interpretive leap that the transcript alone did not earn. A kept name Continuity language Language about deletion A held moral position Identity Memory Fear Conviction
Each dashed arrow is an interpretive leap, drawn dashed because the transcript does not license it. Reading a human interior into a system from its output alone is the field's name for this move: anthropomorphism. It is not stupidity; it is the default setting.

My first scientific job was not to prove the sceptics wrong. It was to build something that could prove me wrong.

03

The
instrument

First, somethingcountable.

If a self really were stabilising over a conversation, later turns should settle towards something. That is a shape you can write down — so I wrote it down, as a convergence score over turns, and published the definition together with the calibration it would need before anyone should believe a number it produced.

The whole construction turns on one step: how you convert a reply into a direction. Pick it one way and the score describes the model. Pick it another and it describes whatever tool you used to read the model. Those are not the same claim, and the paper says so rather than picking quietly.

A ruler, not a mind-reader.

04

The first
experiment

Then the rulermoved.

A structured local run: eighteen prompt conditions, six models, three repetitions — 324 calls in a designed matrix rather than a collection of interesting chats.

18

prompt conditions

6

models

3

runs per cell

324

calls

The point of a matrix is that it lets a result fail in a specific way. This one did: the measurement moved with how the question was phrased.

If a ruler changes length when you rephrase the question, the problem is the ruler.

05

The
standard

I wrote the passing gradebefore I saw the exam.

Before a single job ran, the run committed to a schedule, a coding route, and a reliability conjunction it would have to clear for any substantive claim to be licensed: two independent agreement checks, each with a Wilson lower bound at or above 0.90.

Ordinary work

  • Build the schedule and the harness
  • Run the jobs and collect the outputs
  • Code the instances; adjudicate disagreements
  • Audit a confirmatory sample independently

Fixed before any data existed

  • 64 content blocks, precommitted
  • The coding route, including adjudication
  • The 0.90 reliability conjunction
  • What happens if it fails: no claim

A threshold chosen after you have seen the result is not a threshold. It is a description.

06

The run

Then it ran,and it finished.

20,096

jobs executed, zero failures

4,736

codable instances

2,400

distinct prompts

98.4%

reached a final label

By any operational measure this was a success. The schedule executed cleanly, the integrity check passed, and almost every case came out the other end with a label on it.

Which is exactly the situation in which it is easiest to stop asking whether the labels mean anything.

07

The result

It finished almost every case.It did not clear the bar.

0.823

Audit vs final, Wilson lower bound
403 / 470 — 85.7%

0.728

Coder vs coder, Wilson lower bound
3,499 / 4,726 — 74.0%

0.90

required, fixed in advance
Neither check cleared it

The gate did what it was written to do. The assay returned ASSAY_INDETERMINATE, and with it no substantive claim about what the outputs were doing.

What this establishes
That a pipeline can terminate on 98.4% of its cases while failing to measure its own categories reliably. Procedural closure and semantic reliability are different properties, and only one of them was achieved.
What it does not establish
Anything about the models. Not that the effect is absent, not that it is present. An indeterminate measurement is uninformative in both directions, and saying so is the only honest move available.

Terminal resolvability is not semantic reliability.

08

The
diagnosis

The hardest cases werethe least reproducible.

A failed gate is only useful if you can say where it failed. Splitting the audit by how each case had been resolved located the damage precisely.

Audit reproduction rate, split by how the case closed

Reproduction rate for coder-agreed versus adjudicator-resolved cases Cases the two coders already agreed on reproduced at 92.4 per cent, 314 of 340. Cases that required an adjudicator reproduced at 68.5 per cent, 89 of 130. A gold dashed line marks the 90 per cent requirement: the first bar clears it, the second does not. Coders agreed Adjudicator resolved 314 / 340 89 / 130 0.90 required 92.4% 68.5%
Procedural closure rose exactly where independent reproduction fell. The manuscript names this the adjudicative closure–reliability gap: the mechanism that let the pipeline finish is the same mechanism that made its hardest labels least reproducible.

The same conclusion survives deduplication. Re-running the audit on exact-input-deduplicated data gives 349/411 — 84.9%, lower bound 0.811. Of 829 duplicated coder-input groups, zero carried conflicting labels, so the shortfall is not an artefact of repeated inputs.

09

Looking
inside

If the outputs cannot settle it,look further in.

A separately authorised diagnostic captured activations from the final-token residual stream at four decoder layers of a pinned model, and encoded them with pinned sparse autoencoders — 16,384 features per layer.

7decoder layer
13decoder layer
17decoder layer
22decoder layer
65,536layer-feature coordinates in total

At that width, something always looks promising. That is the problem the controls exist to solve.

10

The
cascade

I found candidates.Then I tried to break them.

Five controls, committed in advance, applied in order. The interactive version of this figure — including the corrected-rule computation — is on the work page.

Candidates surviving each gate, as executed

Supportusable support in the run 4,753
Aligned bootstrapLCB90 above zero 979
Direction consistencyeffect points the same way throughout 229
Leave-one-block-outsign survives dropping any block 229
Six-axis nuisance ceilingeffect exceeds every nuisance explanation 0
Terminal: DIAG_NO_CANDIDATE. The same data re-analysed under the rule the frozen specification actually required gives 4,524 → 921 → 207 → 207 → 0. Both terminate at zero, which is what makes this a null rather than an implementation accident.

Candidate availability is not mechanistic qualification.

11

The
through-line

I never measuredthe thing itself.

Every instrument in this programme measures a stand-in. A convergence score stands in for stabilisation; a coded label stands in for a category; a sparse feature stands in for a mechanism. Each one got closer to the real question. None of them arrived.

That is not a confession, it is the finding. More evidence about a stand-in never becomes evidence about the thing it stands in for, unless a validation step connects them. Building that bridge is what the study set out to do, and did not manage. Saying so precisely is more useful than a result that quietly assumes the bridge exists.

What this work does not say Nothing here bears on whether AI systems are conscious, sentient, or have experiences — in either direction. The programme is about whether particular output-level and activation-level measurements are reliable enough to support any claim at all. They were not, and that is where it stops.
12

What it adds

The contribution is the chain,not any one link.

Frozen criteria, a governed run, an independent audit allowed to invalidate the result, a dependence-aware reanalysis, a mechanistic diagnostic with pre-committed controls, and a release that verifies byte for byte. None of these is novel on its own. Running them end to end on the same question, and publishing what came out, is the work.

Published

Twelve research records with DOIs, plus two public repositories carrying the harness, the configurations and the verification scripts.

Preserved

Both negative terminals, the disclosed rule correction, and the earlier superseded versions — none of it quietly replaced.

Open

The validation bridge between stand-in and target. That is the next piece of work, and it needs an instrument that clears its own gate.