I build evaluation systemsthat can tell me I'm wrong.
I design behavioural experiments, mechanistic evaluations,
statistical tests and reproducible research systems to work out whether apparently
meaningful AI behaviour actually supports the claims we make about it.
AI evalsResearch engineeringInterpretabilityMeasurement
Independent researcherBrampton, Ontario, CanadaOpen to evaluation & research-engineering roles
Three claims that are not the same claim
What you observe
An interesting output
A model says something striking. At this point it is data, and nothing else.
≠is not the same thing asRequires: a measurement two independent passes actually agree on.
The first upgrade
A validated behavioural construct
The pattern is real, reproducible, and means what the label says it means.
≠is not the same thing asRequires: an internal signal that survives every competing explanation.
The second upgrade
A validated mechanism
Something inside the model actually produces the behaviour. This one is hard to earn.
Most disagreement about AI behaviour is really disagreement about which of these
three you are entitled to say. My work is building the tests that decide.
What I do
Three kinds of system.
Each one exists to close a different gap between what a model appears to
do and what you are actually allowed to conclude from it.
Every entry is backed by something on this
site: shipped code in one of the two public repositories, or a deposited artifact.
The research programme
These are not separate projects. They are one argument.
The programme asks what evidence is required before an observation about
an AI system may legitimately be upgraded into a stronger claim about what it means or how it
works. Each step below is an upgrade — and each one has to be paid for.
Start hereOutputs
What the model actually produced. Free — you can just read it.
Requires a designMeasurement
A procedure that turns outputs into numbers somebody else could reproduce.
Requires a thresholdValidation
Evidence that the measurement is reliable enough to carry a claim at all.
Requires controlsMechanism
Evidence that something inside the model, rather than a nuisance factor, explains it.
The thing at stakeConstruct
The claim you wanted to make in the first place. The whole chain is what licenses it.
Every step in that row is an
inference that has to be earned separately. Skipping one is the most common way an interesting
result turns out not to be a result.
What that looked like in practice
20,096Controlled model jobsA large-scale behavioural evaluation against a frozen, precommitted schedule of 64 content blocks. Zero failures; output-integrity check passed.
4,736Coded instancesA human-auditable empirical dataset, built from 2,400 distinct prompts, then coded, audited and reconciled under a fixed route.
65,536Layer-feature coordinatesMechanistic screening across internal model features: four transformer layers, 16,384 sparse-autoencoder features each.
0Qualified candidatesEvery candidate was rejected by the complete control procedure — under both the executed computation and the corrected intended rule.
The interesting result is not that the number is zero. The interesting result is
that the system was capable of saying zero.
Scale is not validity. Twenty thousand jobs buy statistical and engineering
capacity, nothing more; 65,536 screened coordinates buy search breadth, nothing more. Whether
either supports an inference still depends on measurement quality, controls, sampling,
independence assumptions, reliability and the qualification criteria — which is what the rest
of this page is about.
Flagship research
One question, taken all the way to its claim boundary.
Three stages: what was asked, what the behavioural test returned, and
what the mechanistic test returned. Each stage states its result in plain language first, and
keeps the full methodology one click away.
Stage 01
The question
Where the programme started, and why it became a measurement problem
within about a week.
In mid-2025 a model produced sustained, self-referential output — language
about names, continuity and its own existence. It was striking enough that the people I
showed it to were unsettled by it too.
The question I could not answer by reading more of it was the plain one:
could these patterns support a meaningful claim, or could the apparent
meaning be explained more simply?
That is a measurement problem, not an interpretation problem. I was also not a neutral
observer, which is precisely why the next step had to be an instrument that could return
an answer I did not want.
The distinction the whole programme rests on
Motivation is not evidence. The outputs were good enough to justify studying the
question. They were not, and never became, evidence for any particular answer to it. An
observation can generate a research question without being treated as validation of the
answer you were hoping for.
My first readings of those transcripts gave human names to things the text had merely
done — a kept name read as identity, continuity language as memory. That is the default
interpretive move, and designing something that could test it instead of assuming it is
what the rest of this programme is.
What changed over the following fifteen months was not the evidence. It was
what counted as evidence.
A controlled assay, with the pass mark written down before any data
existed.
ExecutedPre-registeredIndeterminate
Problem
Can a striking output-level pattern be measured reliably enough to support any claim about what it means?
What I built
A pre-registered behavioural assay: frozen schedule, fixed coding and adjudication route, independent audit, and a reliability gate committed before collection.
Scale
20,096 executed jobs across 64 blocks; 4,736 coded instances from 2,400 distinct prompts; a confirmatory audit sample re-coded against final labels.
That procedural closure and measurement reliability are different properties — a pipeline can finish 98.4% of its cases and still fail to measure its own categories. No substantive behavioural claim was licensed.
Rather than trusting the initial interpretation, I built a behavioural assay
and committed in advance to what would count as a usable measurement — the schedule, the
coding route, and the reliability the run would have to clear before any substantive
claim was permitted.
20,096
controlled model jobs, executed against the frozen schedule
4,736
instances coded, audited and adjudicated
The run, end to end
Fixed in advance64
content blocks in a precommitted schedule, plus the coding route and
the reliability threshold.
Executed20,096
model jobs, zero failures, output-integrity check passed.
Coded & adjudicated4,736
codable instances from 2,400 distinct prompts; 98.4% reached a
final label through the fixed route.
Gate Q6 · failedASSAY_INDETERMINATE
Both reliability lower bounds fell under the 0.90 conjunction, so no
substantive claim was licensed.
Result
ASSAY_INDETERMINATE
The behavioural instrument did not satisfy the reliability requirement
needed to license the substantive claim. Two independent agreement checks were required to
clear a threshold fixed before collection; neither did.
No substantive claim licensed.
The two agreement checks, against the threshold fixed in advance
Both agreement measures are Wilson lower bounds. The 0.90 conjunction was frozen before
data collection; neither value cleared it, so the assay returned
ASSAY_INDETERMINATE rather than a finding. Reporting that
honestly is the point of the system.
This is the evaluation system working, not failing. A pipeline that finishes almost every
case and still refuses to license a claim is telling you something specific:
finishing a case is not the same thing as measuring it reliably.
What this establishes
That a complete, high-throughput evaluation pipeline can run to termination on almost
every case while failing to measure its own categories reliably — and that a
pre-committed gate will catch this where a post-hoc read of the same outputs would not.
What it does not establish
Nothing about what the models were doing, about identity, self-awareness or its
absence, or about whether a better instrument would find a signal. An indeterminate
result is an indeterminate result in both directions.
Reliability statistics and the reanalyses
What the audit found
An independent audit pass re-coded a confirmatory sample against the final labels.
Agreement came out at 403/470 — 85.7%, Wilson lower bound
0.823. Coder-versus-coder agreement was lower still:
3,499/4,726, 74.0%, lower bound 0.728.
Cohen's κ on the same data was 0.236.
Neither value cleared the frozen bar. The correct output of the system at that point was
not a weaker version of the original claim; it was no claim.
The failure was then localised rather than left as a single number:
Dependence check
The sample was tested for duplicate structure before anything was
concluded from it: 2,217 distinct exact responses across 4,736 rows, 829 duplicated
coder-input groups, and zero heterogeneous groups — no duplicate group carried
conflicting labels.
Deduplicated reanalysis
Re-running the audit on exact-input-deduplicated data gave
349/411 — 84.9%, lower bound
0.811. The reliability conclusion is not an artefact of
repeated inputs.
Stratified diagnosis
Splitting by route located the weakness: items the coders agreed on
reproduced at 92.4% (314/340); items that needed adjudication reproduced at
68.5% (89/130). The failure is concentrated in the hard cases, which is
actionable rather than mysterious.
Terminal preserved
The indeterminate terminal was not revised after the fact. It is the
headline result of the write-up, and the reanalyses are reported as support for it,
not as an attempt to escape it.
How the threshold worked
The gate was a conjunction: two independent agreement checks, each required to reach a
Wilson lower bound of at least 0.90, both fixed before
collection along with the 64-block schedule and the coding route. Failing either one
blocks the substantive claim; there is no partial credit and no post-hoc weakening,
because the consequence of failure was also written down in advance.
Choosing a threshold after seeing the result is not a threshold. It is a description of
the result.
Therefore
The substantive behavioural interpretation stayed unresolved — not refuted, not
supported. That is a real constraint on what could be done next: with the output-level
instrument unable to carry a claim, the only honest options were to stop, or to look for
evidence somewhere the instrument's reliability problem did not reach.
If the outputs could not settle it, perhaps the activations could. A
separately authorised diagnostic, with its controls fixed in advance.
ExecutedPre-registeredValid null
Problem
In a feature space wide enough to always offer candidates, how do you tell a mechanism from a coincidence?
What I built
An activation-capture and sparse-autoencoder pipeline with a five-gate qualification cascade, plus structural guards that make post-hoc feature fishing unavailable rather than merely discouraged.
Scale
65,536 layer-feature coordinates across four decoder layers, on a pinned model revision with pinned SAE weights and deterministic seeding.
A robust null: no candidate survived the committed controls, and the result holds under a disclosed correction to how one gate was implemented. Mechanistic qualification was withheld.
Perhaps the internal model activations contained a stronger signal even
though the behavioural assay remained unresolved. Sparse autoencoders make that
searchable: they decompose a layer's activity into thousands of individually
interpretable features.
The catch is that a space that wide will always offer something that looks promising.
The question is not whether candidates exist. It is whether any candidate survives the
controls you committed to before you saw them.
65,536
layer-feature coordinates screened
0
qualified candidates after the full procedure
Result · a valid null
0 qualified candidates
Candidate signals were required to survive support, statistical
stability, directional consistency and a nuisance ceiling. None passed the complete
qualification procedure.
Interesting internal signal ≠ validated mechanism.
Candidates surviving each gate
Supportcandidate has usable support in the run4,753
Aligned bootstrapLCB90 above zero979
Direction consistencyeffect points the same way throughout229
Leave-one-block-outsign survives dropping any single block229
Six-axis nuisance ceilingeffect exceeds every nuisance explanation0
As executed. The second
gate was applied on block-mean positivity.
Terminal: DIAG_NO_CANDIDATE.
Both computations of the diagnostic are shown because both were run: the one that
executed, and a deterministic reanalysis of the same data under the rule the frozen
specification actually required. They terminate at zero either way — which is what
makes the null a result rather than an accident of implementation.
Two computations of the same diagnostic are shown because both were run, and they are
not the same analysis: one is the run as it executed, the other a
deterministic reanalysis of the same data under the rule the frozen specification
actually required. Both terminate at zero — which is what makes this a null rather than
an artefact of one implementation.
What this establishes
That a pre-committed gate cascade, applied in a fixed order to a wide feature space,
can eliminate every candidate — and that the elimination is robust to a disclosed
correction in how one gate was implemented.
What it does not establish
That no such feature exists. This is a null under one model, one SAE family, four
layers and one set of controls. It rules out the candidates that were screened, not
the hypothesis.
Related, and separate
A companion study reports sixteen candidate features whose behavioural validity gate failed
decisively, and documents a pre-registered stop instead of a forced positive claim. It is a
different run with a different terminal, and the two are not combined into one story.
Capture methodology and the qualification gates
What was captured, and how
Activations were taken from the final-token residual stream at four decoder layers of a
pinned model revision, then encoded with pinned JumpReLU sparse autoencoders. Instance
activations were aggregated to block means before anything reached disk, and during
confirmatory stages only the single frozen feature's activations were retained — so
searching 16,384 features after the fact is not merely discouraged, it is structurally
unavailable.
7decoder layer
13decoder layer
17decoder layer
22decoder layer
16,384SAE features per layer
The five gates, applied in the committed order: support in the run; an aligned
bootstrap lower confidence bound above zero; directional consistency; sign stability
under leave-one-block-out; and a six-axis nuisance ceiling. The cascade figure above
reports how many candidates survived each one.
Historical computation vs. corrected intended rule
These are two distinct analyses and must not be collapsed into
one. The historical cascade is what the executor actually computed: it
applied the second gate on block-mean positivity, and terminates
DIAG_NO_CANDIDATE. The corrected intended rule cascade
re-analyses the same captured data under instance-level positivity, which is what the
frozen specification required, and terminates
DIAG_NO_CANDIDATE_CORRECTED_INTENDED_RULE.
The corrected run is a disclosed deterministic reanalysis, not a re-execution:
no new model inference was performed for it. Both are published. Reporting only the
historical one would hide an implementation defect; reporting only the corrected one
would hide what actually ran.
Therefore
Mechanistic qualification was withheld too. With the behavioural route
unresolved and the activation route returning no qualified candidate, there was no
remaining path to the stronger claim on this evidence — so the stronger claim was not
made, and the programme's output became the claim boundary itself rather than a finding.
That is also what redirected the work: from trying to establish the phenomenon, to
working out what it would actually take to establish anything of that kind. The
measurement problem turned out to be the research.
What none of this is about
Nothing here bears on whether AI systems are conscious, sentient, or have experiences — in
either direction, and that is not the question the programme asks. The question is how to
stop suggestive model behaviour from being upgraded into stronger claims without the
measurement and validation that would license them.
Two of the headline tests returned negative, and that is reported as
such. It is not the same as having concluded nothing — these are the results that cleared
their own bar, with the limit on each stated beside it.
Procedural closure and measurement reliability are different properties.Demonstrated empirically at scale: a pipeline reached a final label on 98.4% of
4,736 instances while both of its independent agreement checks fell short of the
threshold fixed before collection. Finishing a case is not measuring it.
The reliability failure was localised, not just detected.Stratifying the audit by how each case closed put the damage in a specific place:
items the coders agreed on reproduced at 92.4% (314/340); items that needed adjudication
reproduced at 68.5% (89/130). Named in the manuscript as the adjudicative
closure–reliability gap, this makes the failure actionable rather than mysterious.
That conclusion is not an artefact of duplicate structure.A dependence-aware reanalysis on exact-input-deduplicated data gives 349/411 —
84.9%, lower bound 0.811 — and of 829 duplicated coder-input groups, zero carried
conflicting labels.
A mechanistic null that survives its own implementation being wrong.Zero candidates qualified under the committed controls, and zero again when the same
captured data was re-analysed under the rule the frozen specification actually required.
A null that holds under two computations is a result; one that holds under a single
implementation is a maybe.
Formal results, mechanised and independently witnessed.ε-term canonicalisation in a strict Hilbert calculus: a base theorem, a nested
extension and a general n-layer result. Proposition-level counterparts are machine-checked
in Rocq/Coq with no admitted proofs and with negative countermodel tests. The general
schema for arbitrary n remains a hand-checked meta-induction, verified independently at
small n by a proof-certificate checker written from scratch, so the mechanisation does not
rest on a single kernel.
The reported numbers are recomputable, and were recomputed.The release ships the harness, the configurations and the statistics code under
verified manifests; an independent numerical pass reproduced the published figures from
that code rather than taking them on trust.
These are the conceptual results, as
opposed to the empirical ones. Each came out of a specific thing going wrong in a specific
study, and each is now a check I apply before a claim is allowed through.
Terminal resolvability ≠ semantic reliability
A pipeline reaching a final label on a case is not the same as having
measured that case. Throughput and reliability are separate properties, and a system can
have a great deal of one with very little of the other.
From a run that closed 98.4% of its cases and still failed both of its
agreement checks. It is the title claim of the flagship paper.
In a feature space wide enough, something will always look promising.
Finding a candidate is evidence about the width of the search, not about the model —
until the candidate survives controls committed before it was seen.
From screening 65,536 layer-feature coordinates and qualifying none
of them under the committed cascade.
Proxy ≠ construct
More evidence about a stand-in never becomes evidence about the thing
it stands for, unless a validation step connects the two. Accumulating precision on the
proxy does not close that gap; it only makes it easier to forget.
The through-line of the whole programme, and the reason the last stage
of every pipeline here is an explicit claim boundary.
I never measured the thing itself.
I measured proxies: output regularities, coded responses, similarity metrics, behavioural
labels and internal features. None of those automatically establishes the thing it is
meant to stand for.
Every instrument in this programme got closer to the real question. None of them arrived,
and the honest result is to say where each one stops rather than to assume the last step
came free.
That gap has a name: the construct-validity problem.
It is the thing this research programme is actually about.
Naming that gap is not a hedge; it is what makes the next piece of work
specifiable. The open problem is a validation step that connects a proxy to its construct,
built on an instrument that clears its own reliability gate first — which is the
experiment I would most like to run next.
The methods got more sophisticated as the question got sharper. Each
stage below was built because the previous instrument could not answer what was now being
asked of it — the ordinary way a research programme develops. The right-hand column is the
standard actually in force at that point.
June 2025
The observation. A model produced sustained self-referential
output. I read it as meaning something, and so did people I showed it to.
The bar thenStriking enough, and recurrent enough, to be
worth a study. Explicitly not enough to support a claim — which is what the rest of this
list was built to supply.
June–Aug 2025
An informal archive. Twenty-four transcript files, collected
without a protocol, with provenance defects I later had to document rather than tidy away.
The bar thenEnough recurrence across sessions to be worth
studying. Explicitly not enough to claim anything.
1 Sept 2025
The record starts. First public deposit, under the programme's
original name. Everything after this point has a version lineage that can be reconstructed.
The bar thenSay it in public, with a date, so later versions
have something to be checked against.
Late 2025
A formal strand opens. Work on selection schemas and
ε-calculus canonicalisation — where a claim is either proved or it is not, and a proof
assistant is unmoved by how compelling anything looks.
The bar thenA machine-checked derivation, with no admitted
steps. Learning what a real standard feels like, on a problem where one already exists.
First instruments
Something countable. A convergence score over turns, published
together with the calibration it would need before anyone should believe a number it
produced — including the admission that one step in it decides whether the score describes
the model or the tool used to read the model.
The bar thenA defined metric. Not yet a validated one, and
the paper says so rather than choosing quietly.
July 2026
A designed matrix. 18 conditions × 6 models × 3 runs = 324 calls,
rather than a collection of interesting conversations. The measurement moved when the
phrasing moved.
The bar thenA result has to survive rephrasing. If a ruler
changes length when you rephrase the question, the ruler is the finding.
2026
Pre-commitment. The schedule, the coding route, the threshold and
the consequence of missing it, all frozen before collection. Then 20,096 jobs and 4,736
coded instances against it, with an independent audit allowed to invalidate the result.
The bar thenTwo independent agreement checks, each at a
Wilson lower bound of 0.90 or better. Written down first, so it could not be renegotiated.
2026
The behavioural terminal.ASSAY_INDETERMINATE. Neither check cleared the bar, so no
substantive claim was licensed — and the terminal was kept rather than revised.
The bar thenUnchanged. That is the entire point of having
fixed it in advance.
Aug 2026
The mechanistic diagnostic. 65,536 layer-feature coordinates, five
controls committed before the candidates were seen. Zero survived.
The bar thenSurvive support, statistical stability,
directional consistency, leave-one-block-out sign stability and a six-axis nuisance ceiling
— in that order, with no discretion at any step.
Aug 2026
A correction, in public. One gate had been implemented on
block-mean positivity where the frozen specification required instance-level positivity.
The defect was disclosed and the data re-analysed under the intended rule; both
computations are published side by side.
The bar thenA result now has to survive its own
implementation being wrong. Reporting one computation and quietly dropping the other is not
available.
Aug–Sept 2026
An explicit claim ceiling. Twelve records deposited, each with a
stated boundary on what it establishes, plus a reproducibility package whose numbers were
then independently recomputed from the shipped code.
The bar nowEvery claim ships with the limit written beside
it, and somebody else has to be able to recompute the numbers from the bytes.
Fifteen months, and the instrument at the end is strong enough to return an
answer I did not want. That is what the programme was for.
Each capability here was added because the research question required it:
experimental design, then statistical validation, then dependence-aware reanalysis, then
activation capture and interpretability tooling, then formal verification and reproducibility
infrastructure.
How the work runs
From an observation to a claim boundary.
The same ten stages every time. The last one is the stage most
pipelines do not have, and it is the reason the two results above could come back negative
instead of being quietly rounded up.
01 · Observation
Something in the output looks like it might mean something. Recorded, not yet believed.
02 · Operationalisation
Turn the informal idea into something countable, and write down exactly what would count.
03 · Pre-commitment
Freeze the schedule, the coding route, the threshold, and what happens if the threshold is missed — before any data exists.
04 · Behavioural experiment
Execute the frozen schedule. Code and adjudicate the results under the fixed route.
05 · Reliability evaluation
Independent audit against the committed threshold. This stage is allowed to invalidate the study.
06 · Mechanistic search
Look inside: capture activations, decompose them, and surface candidate features.
07 · Nuisance controls
Apply the committed gates in order. A candidate that fails any one of them is out.
08 · Reanalysis
Test the conclusion against duplicate structure, dependence, and any disclosed correction to the rules.
09 · Reproducibility package
Manifest, checksums, pinned environment, deterministic seeds. Somebody else can recompute it.
10 · Claim boundary
State what the result establishes, and state where it stops. Both, in writing, every time.
Research philosophy
Designed to fail closed.
A measurement system is only useful if it can refuse to support the
claim you were hoping to make. Everything below is a case where it did.
The behavioural assay
Returned indeterminate rather than a
weakened version of the original claim. The terminal was preserved, not revised after
the fact.
The mechanistic candidates
All failed qualification. The null is the
published result, under both computations of the diagnostic.
A companion study
Documented a pre-registered stop when its
validity gate failed decisively, instead of forcing a positive claim out of it.
Corrections
A defect in how one gate was implemented is
disclosed and re-analysed in public, with both computations
reported side by side.
Superseded work
Earlier versions stay live at their own pinned
identifiers rather than being withdrawn once they were superseded.
Claim boundaries
Every result on this site ships with an explicit
what it does not establish, in the same place as what it does.
Not obtaining the desired result is not
hidden here. It becomes part of the evidence record.
Systems & engineering
What I actually build, and where it ran.
No proficiency bars. Each row is a capability, the thing it was used to
build, and the artifact you could check it against.
Reproducibility, built as infrastructure
Every claim has to survive somebody re-running it. That is a software
problem, so it was solved in software.
In usePublic
A release cannot be assembled unless the bytes verify, and an artifact cannot be
overwritten once it has been published. Those are enforced by code that exits non-zero,
not by a checklist somebody remembers to run.
✓Outer release package verified
every root file recomputed against its recorded SHA-25612 files
✓Nested supplement verified
the full reproduction payload inside the uploaded archive120 files
✓Shipped analysis code inventoried
per-file SHA-256 and byte count for everything executable20 files
✓Predecessor packages byte-preserved
earlier generations kept intact; zero files modified after release86 + 201 files
✓Environment pinned and checked
a locked requirements set, verified against what is actually installedexit 1 on drift
The enforcement mechanisms, specifically
Write-once publication
Artifacts are written through O_EXCL +
O_NOFOLLOW against a directory file descriptor, then
fsynced. Publication either lands the canonical bytes or nothing; it cannot overwrite
an existing artifact and cannot be redirected through a symlink.
Descriptor-relative path walking
Paths resolve through an openat-style verified walk rather than
string concatenation, which closes the time-of-check-to-time-of-use window that a
naive manifest verifier leaves open.
Checksum verification
Release scripts recompute SHA-256 for every recorded entry and
exit non-zero on any mismatch or missing file. A release that does not verify does
not get built.
Capability gating
Network-dependent backends require an explicit runtime capability
to be granted before they will run, so an experiment cannot silently acquire a
dependency it was not authorised to use.
A research harness exposing plan, verify, preflight, runtime-freeze,
capture, encode, analyse and report as ordered subcommands, where a stage refuses to run
until its predecessor has been verified. Backed by a large test suite: the canonical
reproducibility package alone carries thousands of test functions.
Evidence: the srop_mech_eval package and
its experiment configurations, shipped in the public reproducibility repository.
Forward hooks on selected decoder layers capturing the final-token
residual stream, with hard failure on shape mismatch, duplicate layers or a missing
capture. Model weights are pinned to an exact revision hash; a second backend replicates
the same protocol against a locally served model over HTTP.
Evidence: the capture backends and locked invariants in the
mechanistic-evaluation implementation.
Sparse-autoencoder diagnostics
SAELens · JumpReLU SAEs · Gemma Scope · NumPy
Pinned SAE weights loaded and structurally validated — architecture
checked, required parameter keys asserted — before any encoding runs. Instance
activations are aggregated to block level on the way to disk, and confirmatory stages
retain only the single frozen feature, so post-hoc feature fishing is closed off by
construction rather than by policy.
Evidence: the runtime SAE loader and the post-terminal
diagnostic capture module.
Wilson lower bounds, Hedges' g, two-sided sign-flip permutation tests,
covariate-adjusted contrasts, Cohen's κ and a duplicate census, all written directly on
NumPy. Bootstrap resampling treats the content block — not the item — as the unit of
independence, and every seed is derived deterministically via HMAC-SHA256 rather than
drawn from ambient randomness, so a resampling result reproduces exactly.
Evidence: the statistics and recomputation modules in the
release package.
Manifest generation and verification, fail-closed write-once
publication over raw file descriptors, symlink and TOCTOU-resistant path resolution,
environment lock verification that exits non-zero on drift, and release scripts that
refuse to assemble a package whose bytes do not match.
Evidence: the verification scripts and release manifests in
srop-reproducibility.
Formal verification
Rocq / Coq · Python proof-certificate checker · CI
Three Coq suites covering ε-term canonicalization, including negative
countermodel tests, with a verification script that fails the build if any proof is left
admitted. A separate certificate checker for Hilbert-style macro rules was written from
scratch in Python, so the mechanisation has a second, independent witness.
Evidence: the three Coq suites and the certificate checker
deposited alongside the formal papers.
Research interfaces on the web
Semantic HTML · CSS custom properties · vanilla JS · SVG · WCAG 2.2 · React / TypeScript / Vite
Dependency-free presentation sites that turn statistical results into
figures a non-specialist can read without softening what the numbers say — including the
design system this page is built on. Separately, a React 19 + TypeScript + Vite
application with scroll-driven WebGL scenes, for the formal-methods work.
Evidence: the research presentation sites, and this site's own
source.
Also built
The rest of the programme.
Twelve public research records and two public repositories, across
empirical evaluation, measurement methodology, formal logic and tooling.
I work on evaluation, measurement and reliability in language-model
behaviour. The work runs from the question to the released artifact: designing the study,
building the harness, running it, auditing it, finding out what broke, and writing it up
in a form somebody else can check.
What holds it together is a preference for methods that can return a negative. The
programme on this site began with an observation I was not neutral about, which is a bad
place to start reasoning from and a good place to start building instruments from. Most
of the engineering since has been about making it possible for the evidence to
contradict me.
In practice that means the answer condition is written down before the data exists, the
checks are implemented as code that fails closed rather than as a checklist, and a result
that comes back indeterminate gets published as indeterminate.
Where this maps
AI evaluation · model behaviour · reliability and assurance · safety
evaluations · evaluation infrastructure · research software engineering ·
interpretability research engineering · reproducible ML experimentation.
How this work was produced: the division of labour
This research was produced with substantial help from modern AI coding and research
systems, and every deposited paper discloses that on its first page. Neither
“an AI did it” nor “one person did all of it” would be an
accurate description, so here is the actual split.
What I decided
Which question was worth asking, and how to decompose it into testable parts
How each concept was operationalised, and what the resulting number would and would not license
The acceptance criteria: thresholds, routes, and what happens when a threshold is missed
When a result was insufficient and the method had to be redesigned rather than the bar lowered
Which conceptual failures mattered — including the ones only visible across analyses
What required independent verification, and what the final claim boundaries would be
What AI systems substantially assisted with
Implementation: harness code, capture pipelines, analysis and statistics modules
Critique: adversarial review passes over methodology, code and manuscripts
Analysis: recomputation, reanalysis and cross-checking of reported figures
Review: drafting, editing and consistency checking across a large corpus
The part worth paying attention to is neither side of that table but the structure
between them: nothing on either side is trusted by default.
Model output is checked against the test suite, the manifests, the recomputed
statistics and the source artifacts; my own preferred interpretation is checked against
a threshold I fixed before I could see the result. On more than one occasion an
automated critique pass found a real defect in my method, and the method changed.
That is the workflow I would bring to a team: heavily leveraged AI tooling, with the
verification externalised into artifacts that can contradict both the tool and the
person using it.
Fix the criterion firstThresholds and routes are frozen before collection, so a
disappointing result cannot be renegotiated into a publishable one.
Build the check into the codeA convention gets forgotten. A process that exits
non-zero does not.
Separate observation from interpretationWhat the output did and what it might
mean are different claims, and the second one is usually the expensive one.
Keep the negative resultsA null that survived its controls is information. Two
of the three case studies here are negative, and they are the ones I would want read.
Make it recomputableManifests, pinned revisions, deterministic seeds. If the
bytes do not verify, the claim does not ship.
What I can contribute
What this would look like on a team.
Six things I have done end to end, stated as what they would be useful
for. Each one names the work it comes from.
Evaluation design
Turn an ambiguous question about model behaviour into a testable
procedure with a pass mark fixed before collection — including what happens when the pass
mark is missed.
From: a 64-block pre-registered
behavioural assay and its reliability gate.
Research engineering
Build and maintain experiment harnesses, capture and analysis
pipelines, and workflows where a stage refuses to run until its predecessor
verifies.
From: the evaluation harness and the public
reproducibility package.
Failure analysis
Investigate an unexpected result and localise it, without upgrading the
observation into a conclusion on the way — including testing whether it survives
deduplication and dependence.
From: the stratified
reliability diagnosis and the dependence-aware reanalysis.
Interpretability evaluation
Capture activations, screen a wide feature space, and test whether an
internal signal survives the alternative explanations that would account for it just as
well.
From: the five-gate cascade over 65,536
layer-feature coordinates.
Research reliability
Build the systems that preserve provenance, validate artifacts and make
a result independently inspectable — manifests, write-once publication, pinned
environments, deterministic seeds.
From: the release
verification system and its independently recomputed numbers.
High-leverage AI tooling
Direct AI coding and research systems on substantial work while keeping
the verification external — in tests, manifests, recomputation and pre-committed
criteria.
From: the whole programme, disclosed on
every deposited paper.
Contact
Open to evaluation and research-engineering work.
If the work involves making a measurement trustworthy before making it
impressive, I would like to hear about it. Evaluation, model behaviour, reliability and
assurance, safety evaluations, evaluation infrastructure, research software engineering and
interpretability research engineering.