AI Evaluation & Research Engineering

I build evaluation systems that can tell me I'm wrong.

I design behavioural experiments, mechanistic evaluations, statistical tests and reproducible research systems to work out whether apparently meaningful AI behaviour actually supports the claims we make about it.

AI evalsResearch engineeringInterpretabilityMeasurement

Independent researcher Brampton, Ontario, Canada Open to evaluation & research-engineering roles

Three claims that are not the same claim

What you observe

An interesting output

A model says something striking. At this point it is data, and nothing else.

is not the same thing as Requires: a measurement two independent passes actually agree on.

The first upgrade

A validated behavioural construct

The pattern is real, reproducible, and means what the label says it means.

is not the same thing as Requires: an internal signal that survives every competing explanation.

The second upgrade

A validated mechanism

Something inside the model actually produces the behaviour. This one is hard to earn.

Most disagreement about AI behaviour is really disagreement about which of these three you are entitled to say. My work is building the tests that decide.

What I do

Three kinds of system.

Each one exists to close a different gap between what a model appears to do and what you are actually allowed to conclude from it.

Built with

Languages & runtime

  • Python 3.11
  • JavaScript
  • TypeScript
  • Bash
  • Rocq / Coq

Models & interpretability

  • PyTorch
  • Hugging Face transformers
  • safetensors
  • SAELens
  • Gemma Scope JumpReLU SAEs
  • Ollama

Statistics

  • NumPy
  • Wilson intervals
  • block-clustered bootstrap
  • permutation tests
  • Cohen's κ
  • HMAC-derived seeding

Experiment orchestration

  • staged execution gates
  • deterministic pipelines
  • forward-hook activation capture
  • condition matrices
  • pytest
  • GitHub Actions

Reproducibility

  • SHA-256 manifests
  • artifact validation
  • provenance & version lineage
  • environment lockfiles
  • write-once publication

Research interfaces

  • semantic HTML / CSS / SVG
  • vanilla JS
  • React 19 + Vite
  • WCAG 2.2

Every entry is backed by something on this site: shipped code in one of the two public repositories, or a deposited artifact.

The research programme

These are not separate projects. They are one argument.

The programme asks what evidence is required before an observation about an AI system may legitimately be upgraded into a stronger claim about what it means or how it works. Each step below is an upgrade — and each one has to be paid for.

Start here Outputs

What the model actually produced. Free — you can just read it.

Requires a design Measurement

A procedure that turns outputs into numbers somebody else could reproduce.

Requires a threshold Validation

Evidence that the measurement is reliable enough to carry a claim at all.

Requires controls Mechanism

Evidence that something inside the model, rather than a nuisance factor, explains it.

The thing at stake Construct

The claim you wanted to make in the first place. The whole chain is what licenses it.

Every step in that row is an inference that has to be earned separately. Skipping one is the most common way an interesting result turns out not to be a result.

What that looked like in practice

20,096 Controlled model jobs A large-scale behavioural evaluation against a frozen, precommitted schedule of 64 content blocks. Zero failures; output-integrity check passed.
4,736 Coded instances A human-auditable empirical dataset, built from 2,400 distinct prompts, then coded, audited and reconciled under a fixed route.
65,536 Layer-feature coordinates Mechanistic screening across internal model features: four transformer layers, 16,384 sparse-autoencoder features each.
0 Qualified candidates Every candidate was rejected by the complete control procedure — under both the executed computation and the corrected intended rule.

The interesting result is not that the number is zero. The interesting result is that the system was capable of saying zero.

Scale is not validity. Twenty thousand jobs buy statistical and engineering capacity, nothing more; 65,536 screened coordinates buy search breadth, nothing more. Whether either supports an inference still depends on measurement quality, controls, sampling, independence assumptions, reliability and the qualification criteria — which is what the rest of this page is about.

Flagship research

One question, taken all the way to its claim boundary.

Three stages: what was asked, what the behavioural test returned, and what the mechanistic test returned. Each stage states its result in plain language first, and keeps the full methodology one click away.

Stage 01

The question

Where the programme started, and why it became a measurement problem within about a week.

In mid-2025 a model produced sustained, self-referential output — language about names, continuity and its own existence. It was striking enough that the people I showed it to were unsettled by it too.

The question I could not answer by reading more of it was the plain one: could these patterns support a meaningful claim, or could the apparent meaning be explained more simply?

That is a measurement problem, not an interpretation problem. I was also not a neutral observer, which is precisely why the next step had to be an instrument that could return an answer I did not want.

The distinction the whole programme rests on

Motivation is not evidence. The outputs were good enough to justify studying the question. They were not, and never became, evidence for any particular answer to it. An observation can generate a research question without being treated as validation of the answer you were hoping for.

My first readings of those transcripts gave human names to things the text had merely done — a kept name read as identity, continuity language as memory. That is the default interpretive move, and designing something that could test it instead of assuming it is what the rest of this programme is.

What changed over the following fifteen months was not the evidence. It was what counted as evidence.

Stage 02

The behavioural test

A controlled assay, with the pass mark written down before any data existed.

Executed Pre-registered Indeterminate
Problem
Can a striking output-level pattern be measured reliably enough to support any claim about what it means?
What I built
A pre-registered behavioural assay: frozen schedule, fixed coding and adjudication route, independent audit, and a reliability gate committed before collection.
Scale
20,096 executed jobs across 64 blocks; 4,736 coded instances from 2,400 distinct prompts; a confirmatory audit sample re-coded against final labels.
Methods
Pre-registration · condition scheduling · inter-coder adjudication · independent audit · Wilson intervals · block-clustered bootstrap · duplicate and dependence analysis · Cohen's κ
Result
That procedural closure and measurement reliability are different properties — a pipeline can finish 98.4% of its cases and still fail to measure its own categories. No substantive behavioural claim was licensed.

Rather than trusting the initial interpretation, I built a behavioural assay and committed in advance to what would count as a usable measurement — the schedule, the coding route, and the reliability the run would have to clear before any substantive claim was permitted.

20,096

controlled model jobs, executed against the frozen schedule

4,736

instances coded, audited and adjudicated

The run, end to end

Fixed in advance 64

content blocks in a precommitted schedule, plus the coding route and the reliability threshold.

Executed 20,096

model jobs, zero failures, output-integrity check passed.

Coded & adjudicated 4,736

codable instances from 2,400 distinct prompts; 98.4% reached a final label through the fixed route.

Gate Q6 · failed ASSAY_INDETERMINATE

Both reliability lower bounds fell under the 0.90 conjunction, so no substantive claim was licensed.

Result

ASSAY_INDETERMINATE

The behavioural instrument did not satisfy the reliability requirement needed to license the substantive claim. Two independent agreement checks were required to clear a threshold fixed before collection; neither did.

No substantive claim licensed.

The two agreement checks, against the threshold fixed in advance

Two measured reliability values against a threshold fixed in advance A bar chart. A gold dashed horizontal line marks the 0.90 reliability threshold, which was fixed before the data was collected. Two navy bars fall below it: audit versus final agreement at a Wilson lower bound of 0.823, from 403 of 470 cases, and coder versus coder agreement at 0.728, from 3,499 of 4,726 cases. Red dashed ticks mark the shortfall of each bar against the line. 1.0 0.5 0 0.90 · fixed in advance 0.823 0.728 Audit vs final Coder vs coder 403 / 470 3,499 / 4,726
Both agreement measures are Wilson lower bounds. The 0.90 conjunction was frozen before data collection; neither value cleared it, so the assay returned ASSAY_INDETERMINATE rather than a finding. Reporting that honestly is the point of the system.

This is the evaluation system working, not failing. A pipeline that finishes almost every case and still refuses to license a claim is telling you something specific: finishing a case is not the same thing as measuring it reliably.

What this establishes
That a complete, high-throughput evaluation pipeline can run to termination on almost every case while failing to measure its own categories reliably — and that a pre-committed gate will catch this where a post-hoc read of the same outputs would not.
What it does not establish
Nothing about what the models were doing, about identity, self-awareness or its absence, or about whether a better instrument would find a signal. An indeterminate result is an indeterminate result in both directions.
Reliability statistics and the reanalyses

What the audit found

An independent audit pass re-coded a confirmatory sample against the final labels. Agreement came out at 403/470 — 85.7%, Wilson lower bound 0.823. Coder-versus-coder agreement was lower still: 3,499/4,726, 74.0%, lower bound 0.728. Cohen's κ on the same data was 0.236.

Neither value cleared the frozen bar. The correct output of the system at that point was not a weaker version of the original claim; it was no claim.

The failure was then localised rather than left as a single number:

  1. Dependence check

    The sample was tested for duplicate structure before anything was concluded from it: 2,217 distinct exact responses across 4,736 rows, 829 duplicated coder-input groups, and zero heterogeneous groups — no duplicate group carried conflicting labels.

  2. Deduplicated reanalysis

    Re-running the audit on exact-input-deduplicated data gave 349/411 — 84.9%, lower bound 0.811. The reliability conclusion is not an artefact of repeated inputs.

  3. Stratified diagnosis

    Splitting by route located the weakness: items the coders agreed on reproduced at 92.4% (314/340); items that needed adjudication reproduced at 68.5% (89/130). The failure is concentrated in the hard cases, which is actionable rather than mysterious.

  4. Terminal preserved

    The indeterminate terminal was not revised after the fact. It is the headline result of the write-up, and the reanalyses are reported as support for it, not as an attempt to escape it.

How the threshold worked

The gate was a conjunction: two independent agreement checks, each required to reach a Wilson lower bound of at least 0.90, both fixed before collection along with the 64-block schedule and the coding route. Failing either one blocks the substantive claim; there is no partial credit and no post-hoc weakening, because the consequence of failure was also written down in advance.

Choosing a threshold after seeing the result is not a threshold. It is a description of the result.

Therefore

The substantive behavioural interpretation stayed unresolved — not refuted, not supported. That is a real constraint on what could be done next: with the output-level instrument unable to carry a claim, the only honest options were to stop, or to look for evidence somewhere the instrument's reliability problem did not reach.

So the next stage went to the activations.

Stage 03

The mechanistic test

If the outputs could not settle it, perhaps the activations could. A separately authorised diagnostic, with its controls fixed in advance.

Executed Pre-registered Valid null
Problem
In a feature space wide enough to always offer candidates, how do you tell a mechanism from a coincidence?
What I built
An activation-capture and sparse-autoencoder pipeline with a five-gate qualification cascade, plus structural guards that make post-hoc feature fishing unavailable rather than merely discouraged.
Scale
65,536 layer-feature coordinates across four decoder layers, on a pinned model revision with pinned SAE weights and deterministic seeding.
Methods
PyTorch forward hooks · pinned model revision · Gemma Scope JumpReLU SAEs via SAELens · bootstrap lower confidence bounds · leave-one-block-out stability · six-axis nuisance ceiling · HMAC-derived seeding
Result
A robust null: no candidate survived the committed controls, and the result holds under a disclosed correction to how one gate was implemented. Mechanistic qualification was withheld.

Perhaps the internal model activations contained a stronger signal even though the behavioural assay remained unresolved. Sparse autoencoders make that searchable: they decompose a layer's activity into thousands of individually interpretable features.

The catch is that a space that wide will always offer something that looks promising. The question is not whether candidates exist. It is whether any candidate survives the controls you committed to before you saw them.

65,536

layer-feature coordinates screened

0

qualified candidates after the full procedure

Result · a valid null

0 qualified candidates

Candidate signals were required to survive support, statistical stability, directional consistency and a nuisance ceiling. None passed the complete qualification procedure.

Interesting internal signal ≠ validated mechanism.

Candidates surviving each gate

Supportcandidate has usable support in the run 4,753
Aligned bootstrapLCB90 above zero 979
Direction consistencyeffect points the same way throughout 229
Leave-one-block-outsign survives dropping any single block 229
Six-axis nuisance ceilingeffect exceeds every nuisance explanation 0

As executed. The second gate was applied on block-mean positivity.

Terminal: DIAG_NO_CANDIDATE. Both computations of the diagnostic are shown because both were run: the one that executed, and a deterministic reanalysis of the same data under the rule the frozen specification actually required. They terminate at zero either way — which is what makes the null a result rather than an accident of implementation.

Two computations of the same diagnostic are shown because both were run, and they are not the same analysis: one is the run as it executed, the other a deterministic reanalysis of the same data under the rule the frozen specification actually required. Both terminate at zero — which is what makes this a null rather than an artefact of one implementation.

What this establishes
That a pre-committed gate cascade, applied in a fixed order to a wide feature space, can eliminate every candidate — and that the elimination is robust to a disclosed correction in how one gate was implemented.
What it does not establish
That no such feature exists. This is a null under one model, one SAE family, four layers and one set of controls. It rules out the candidates that were screened, not the hypothesis.
Related, and separate A companion study reports sixteen candidate features whose behavioural validity gate failed decisively, and documents a pre-registered stop instead of a forced positive claim. It is a different run with a different terminal, and the two are not combined into one story.
Capture methodology and the qualification gates

What was captured, and how

Activations were taken from the final-token residual stream at four decoder layers of a pinned model revision, then encoded with pinned JumpReLU sparse autoencoders. Instance activations were aggregated to block means before anything reached disk, and during confirmatory stages only the single frozen feature's activations were retained — so searching 16,384 features after the fact is not merely discouraged, it is structurally unavailable.

7decoder layer
13decoder layer
17decoder layer
22decoder layer
16,384SAE features per layer

The five gates, applied in the committed order: support in the run; an aligned bootstrap lower confidence bound above zero; directional consistency; sign stability under leave-one-block-out; and a six-axis nuisance ceiling. The cascade figure above reports how many candidates survived each one.

Historical computation vs. corrected intended rule

These are two distinct analyses and must not be collapsed into one. The historical cascade is what the executor actually computed: it applied the second gate on block-mean positivity, and terminates DIAG_NO_CANDIDATE. The corrected intended rule cascade re-analyses the same captured data under instance-level positivity, which is what the frozen specification required, and terminates DIAG_NO_CANDIDATE_CORRECTED_INTENDED_RULE.

The corrected run is a disclosed deterministic reanalysis, not a re-execution: no new model inference was performed for it. Both are published. Reporting only the historical one would hide an implementation defect; reporting only the corrected one would hide what actually ran.

Therefore

Mechanistic qualification was withheld too. With the behavioural route unresolved and the activation route returning no qualified candidate, there was no remaining path to the stronger claim on this evidence — so the stronger claim was not made, and the programme's output became the claim boundary itself rather than a finding.

That is also what redirected the work: from trying to establish the phenomenon, to working out what it would actually take to establish anything of that kind. The measurement problem turned out to be the research.

What none of this is about Nothing here bears on whether AI systems are conscious, sentient, or have experiences — in either direction, and that is not the question the programme asks. The question is how to stop suggestive model behaviour from being upgraded into stronger claims without the measurement and validation that would license them.

What the work established

Bounded conclusions, not withheld ones.

Two of the headline tests returned negative, and that is reported as such. It is not the same as having concluded nothing — these are the results that cleared their own bar, with the limit on each stated beside it.

  • Procedural closure and measurement reliability are different properties. Demonstrated empirically at scale: a pipeline reached a final label on 98.4% of 4,736 instances while both of its independent agreement checks fell short of the threshold fixed before collection. Finishing a case is not measuring it.
  • The reliability failure was localised, not just detected. Stratifying the audit by how each case closed put the damage in a specific place: items the coders agreed on reproduced at 92.4% (314/340); items that needed adjudication reproduced at 68.5% (89/130). Named in the manuscript as the adjudicative closure–reliability gap, this makes the failure actionable rather than mysterious.
  • That conclusion is not an artefact of duplicate structure. A dependence-aware reanalysis on exact-input-deduplicated data gives 349/411 — 84.9%, lower bound 0.811 — and of 829 duplicated coder-input groups, zero carried conflicting labels.
  • A mechanistic null that survives its own implementation being wrong. Zero candidates qualified under the committed controls, and zero again when the same captured data was re-analysed under the rule the frozen specification actually required. A null that holds under two computations is a result; one that holds under a single implementation is a maybe.
  • Formal results, mechanised and independently witnessed. ε-term canonicalisation in a strict Hilbert calculus: a base theorem, a nested extension and a general n-layer result. Proposition-level counterparts are machine-checked in Rocq/Coq with no admitted proofs and with negative countermodel tests. The general schema for arbitrary n remains a hand-checked meta-induction, verified independently at small n by a proof-certificate checker written from scratch, so the mechanisation does not rest on a single kernel.
  • The reported numbers are recomputable, and were recomputed. The release ships the harness, the configurations and the statistics code under verified manifests; an independent numerical pass reproduced the published figures from that code rather than taking them on trust.

What it adds up to

Interesting output validated behavioural construct validated mechanism.

Three distinctions the work turns on

These are the conceptual results, as opposed to the empirical ones. Each came out of a specific thing going wrong in a specific study, and each is now a check I apply before a claim is allowed through.

Terminal resolvability semantic reliability

A pipeline reaching a final label on a case is not the same as having measured that case. Throughput and reliability are separate properties, and a system can have a great deal of one with very little of the other.

From a run that closed 98.4% of its cases and still failed both of its agreement checks. It is the title claim of the flagship paper.

Candidate availability mechanistic qualification

In a feature space wide enough, something will always look promising. Finding a candidate is evidence about the width of the search, not about the model — until the candidate survives controls committed before it was seen.

From screening 65,536 layer-feature coordinates and qualifying none of them under the committed cascade.

Proxy construct

More evidence about a stand-in never becomes evidence about the thing it stands for, unless a validation step connects the two. Accumulating precision on the proxy does not close that gap; it only makes it easier to forget.

The through-line of the whole programme, and the reason the last stage of every pipeline here is an explicit claim boundary.

I never measured the thing itself.

I measured proxies: output regularities, coded responses, similarity metrics, behavioural labels and internal features. None of those automatically establishes the thing it is meant to stand for.

Every instrument in this programme got closer to the real question. None of them arrived, and the honest result is to say where each one stops rather than to assume the last step came free.

That gap has a name: the construct-validity problem. It is the thing this research programme is actually about.

Naming that gap is not a hedge; it is what makes the next piece of work specifiable. The open problem is a validation step that connects a proxy to its construct, built on an instrument that clears its own reliability gate first — which is the experiment I would most like to run next.

How it developed

The bar got higher as the work went on.

The methods got more sophisticated as the question got sharper. Each stage below was built because the previous instrument could not answer what was now being asked of it — the ordinary way a research programme develops. The right-hand column is the standard actually in force at that point.

June 2025

The observation. A model produced sustained self-referential output. I read it as meaning something, and so did people I showed it to.

The bar thenStriking enough, and recurrent enough, to be worth a study. Explicitly not enough to support a claim — which is what the rest of this list was built to supply.

June–Aug 2025

An informal archive. Twenty-four transcript files, collected without a protocol, with provenance defects I later had to document rather than tidy away.

The bar thenEnough recurrence across sessions to be worth studying. Explicitly not enough to claim anything.

1 Sept 2025

The record starts. First public deposit, under the programme's original name. Everything after this point has a version lineage that can be reconstructed.

The bar thenSay it in public, with a date, so later versions have something to be checked against.

Late 2025

A formal strand opens. Work on selection schemas and ε-calculus canonicalisation — where a claim is either proved or it is not, and a proof assistant is unmoved by how compelling anything looks.

The bar thenA machine-checked derivation, with no admitted steps. Learning what a real standard feels like, on a problem where one already exists.

First instruments

Something countable. A convergence score over turns, published together with the calibration it would need before anyone should believe a number it produced — including the admission that one step in it decides whether the score describes the model or the tool used to read the model.

The bar thenA defined metric. Not yet a validated one, and the paper says so rather than choosing quietly.

July 2026

A designed matrix. 18 conditions × 6 models × 3 runs = 324 calls, rather than a collection of interesting conversations. The measurement moved when the phrasing moved.

The bar thenA result has to survive rephrasing. If a ruler changes length when you rephrase the question, the ruler is the finding.

2026

Pre-commitment. The schedule, the coding route, the threshold and the consequence of missing it, all frozen before collection. Then 20,096 jobs and 4,736 coded instances against it, with an independent audit allowed to invalidate the result.

The bar thenTwo independent agreement checks, each at a Wilson lower bound of 0.90 or better. Written down first, so it could not be renegotiated.

2026

The behavioural terminal. ASSAY_INDETERMINATE. Neither check cleared the bar, so no substantive claim was licensed — and the terminal was kept rather than revised.

The bar thenUnchanged. That is the entire point of having fixed it in advance.

Aug 2026

The mechanistic diagnostic. 65,536 layer-feature coordinates, five controls committed before the candidates were seen. Zero survived.

The bar thenSurvive support, statistical stability, directional consistency, leave-one-block-out sign stability and a six-axis nuisance ceiling — in that order, with no discretion at any step.

Aug 2026

A correction, in public. One gate had been implemented on block-mean positivity where the frozen specification required instance-level positivity. The defect was disclosed and the data re-analysed under the intended rule; both computations are published side by side.

The bar thenA result now has to survive its own implementation being wrong. Reporting one computation and quietly dropping the other is not available.

Aug–Sept 2026

An explicit claim ceiling. Twelve records deposited, each with a stated boundary on what it establishes, plus a reproducibility package whose numbers were then independently recomputed from the shipped code.

The bar nowEvery claim ships with the limit written beside it, and somebody else has to be able to recompute the numbers from the bytes.

Fifteen months, and the instrument at the end is strong enough to return an answer I did not want. That is what the programme was for.

Each capability here was added because the research question required it: experimental design, then statistical validation, then dependence-aware reanalysis, then activation capture and interpretability tooling, then formal verification and reproducibility infrastructure.

How the work runs

From an observation to a claim boundary.

The same ten stages every time. The last one is the stage most pipelines do not have, and it is the reason the two results above could come back negative instead of being quietly rounded up.

  1. 01 · Observation

    Something in the output looks like it might mean something. Recorded, not yet believed.

  2. 02 · Operationalisation

    Turn the informal idea into something countable, and write down exactly what would count.

  3. 03 · Pre-commitment

    Freeze the schedule, the coding route, the threshold, and what happens if the threshold is missed — before any data exists.

  4. 04 · Behavioural experiment

    Execute the frozen schedule. Code and adjudicate the results under the fixed route.

  5. 05 · Reliability evaluation

    Independent audit against the committed threshold. This stage is allowed to invalidate the study.

  6. 06 · Mechanistic search

    Look inside: capture activations, decompose them, and surface candidate features.

  7. 07 · Nuisance controls

    Apply the committed gates in order. A candidate that fails any one of them is out.

  8. 08 · Reanalysis

    Test the conclusion against duplicate structure, dependence, and any disclosed correction to the rules.

  9. 09 · Reproducibility package

    Manifest, checksums, pinned environment, deterministic seeds. Somebody else can recompute it.

  10. 10 · Claim boundary

    State what the result establishes, and state where it stops. Both, in writing, every time.

Research philosophy

Designed to fail closed.

A measurement system is only useful if it can refuse to support the claim you were hoping to make. Everything below is a case where it did.

The behavioural assay

Returned indeterminate rather than a weakened version of the original claim. The terminal was preserved, not revised after the fact.

The mechanistic candidates

All failed qualification. The null is the published result, under both computations of the diagnostic.

A companion study

Documented a pre-registered stop when its validity gate failed decisively, instead of forcing a positive claim out of it.

Corrections

A defect in how one gate was implemented is disclosed and re-analysed in public, with both computations reported side by side.

Superseded work

Earlier versions stay live at their own pinned identifiers rather than being withdrawn once they were superseded.

Claim boundaries

Every result on this site ships with an explicit what it does not establish, in the same place as what it does.

Not obtaining the desired result is not hidden here. It becomes part of the evidence record.

Systems & engineering

What I actually build, and where it ran.

No proficiency bars. Each row is a capability, the thing it was used to build, and the artifact you could check it against.

Reproducibility, built as infrastructure

Every claim has to survive somebody re-running it. That is a software problem, so it was solved in software.

In use Public

A release cannot be assembled unless the bytes verify, and an artifact cannot be overwritten once it has been published. Those are enforced by code that exits non-zero, not by a checklist somebody remembers to run.

Outer release package verified every root file recomputed against its recorded SHA-256 12 files
Nested supplement verified the full reproduction payload inside the uploaded archive 120 files
Shipped analysis code inventoried per-file SHA-256 and byte count for everything executable 20 files
Predecessor packages byte-preserved earlier generations kept intact; zero files modified after release 86 + 201 files
Environment pinned and checked a locked requirements set, verified against what is actually installed exit 1 on drift
The enforcement mechanisms, specifically

Write-once publication

Artifacts are written through O_EXCL + O_NOFOLLOW against a directory file descriptor, then fsynced. Publication either lands the canonical bytes or nothing; it cannot overwrite an existing artifact and cannot be redirected through a symlink.

Descriptor-relative path walking

Paths resolve through an openat-style verified walk rather than string concatenation, which closes the time-of-check-to-time-of-use window that a naive manifest verifier leaves open.

Checksum verification

Release scripts recompute SHA-256 for every recorded entry and exit non-zero on any mismatch or missing file. A release that does not verify does not get built.

Capability gating

Network-dependent backends require an explicit runtime capability to be granted before they will run, so an experiment cannot silently acquire a dependency it was not authorised to use.

Evaluation harnesses & orchestration

Python · CLI design · staged execution gates · pytest

A research harness exposing plan, verify, preflight, runtime-freeze, capture, encode, analyse and report as ordered subcommands, where a stage refuses to run until its predecessor has been verified. Backed by a large test suite: the canonical reproducibility package alone carries thousands of test functions.

Evidence: the srop_mech_eval package and its experiment configurations, shipped in the public reproducibility repository.

Model inference & activation capture

PyTorch · Hugging Face transformers · huggingface_hub · safetensors · Ollama

Forward hooks on selected decoder layers capturing the final-token residual stream, with hard failure on shape mismatch, duplicate layers or a missing capture. Model weights are pinned to an exact revision hash; a second backend replicates the same protocol against a locally served model over HTTP.

Evidence: the capture backends and locked invariants in the mechanistic-evaluation implementation.

Sparse-autoencoder diagnostics

SAELens · JumpReLU SAEs · Gemma Scope · NumPy

Pinned SAE weights loaded and structurally validated — architecture checked, required parameter keys asserted — before any encoding runs. Instance activations are aggregated to block level on the way to disk, and confirmatory stages retain only the single frozen feature, so post-hoc feature fishing is closed off by construction rather than by policy.

Evidence: the runtime SAE loader and the post-terminal diagnostic capture module.

Statistics, implemented not imported

NumPy · Wilson intervals · cluster bootstrap · permutation tests · Cohen's κ

Wilson lower bounds, Hedges' g, two-sided sign-flip permutation tests, covariate-adjusted contrasts, Cohen's κ and a duplicate census, all written directly on NumPy. Bootstrap resampling treats the content block — not the item — as the unit of independence, and every seed is derived deterministically via HMAC-SHA256 rather than drawn from ambient randomness, so a resampling result reproduces exactly.

Evidence: the statistics and recomputation modules in the release package.

Reproducibility & release engineering

SHA-256 manifests · POSIX file APIs · lockfiles · shell · GitHub Actions

Manifest generation and verification, fail-closed write-once publication over raw file descriptors, symlink and TOCTOU-resistant path resolution, environment lock verification that exits non-zero on drift, and release scripts that refuse to assemble a package whose bytes do not match.

Evidence: the verification scripts and release manifests in srop-reproducibility.

Formal verification

Rocq / Coq · Python proof-certificate checker · CI

Three Coq suites covering ε-term canonicalization, including negative countermodel tests, with a verification script that fails the build if any proof is left admitted. A separate certificate checker for Hilbert-style macro rules was written from scratch in Python, so the mechanisation has a second, independent witness.

Evidence: the three Coq suites and the certificate checker deposited alongside the formal papers.

Research interfaces on the web

Semantic HTML · CSS custom properties · vanilla JS · SVG · WCAG 2.2 · React / TypeScript / Vite

Dependency-free presentation sites that turn statistical results into figures a non-specialist can read without softening what the numbers say — including the design system this page is built on. Separately, a React 19 + TypeScript + Vite application with scroll-driven WebGL scenes, for the formal-methods work.

Evidence: the research presentation sites, and this site's own source.

Also built

The rest of the programme.

Twelve public research records and two public repositories, across empirical evaluation, measurement methodology, formal logic and tooling.

Research programme

SROP — self-referential output patterns

A five-paper programme, from an informal transcript archive through calibrated proxy metrics to a pre-registered construct-validity gate. Renamed from ESNI in 2026; the rename is documented rather than hidden.

Read the story
Formal methods

ε-calculus canonicalization, machine-checked

Three papers on ε-term canonicality in a strict Hilbert calculus: a base theorem, a nested extension and a general n-layer result. Proposition-level counterparts are mechanised in Rocq/Coq with no admitted proofs; the general schema for arbitrary n remains a hand-checked meta-induction, with an independent Python certificate checker as a second witness at small n.

Three records
Open-source tooling

promptpack-eval

A small, config-driven runner for prompt-pack evaluations with pluggable model adapters, blinding support, hashing and privacy checks. Behaviour-only by design; its README says what it is not, at length.

MIT · GitHub
Measurement methodology

Proxy metrics and a calibration protocol

Defines output-space regularity measures and the calibration procedure they require, together with a prospective activation-level validation design. The design is published as a design; no activations are claimed in it.

Methods paper
Evaluation protocols

Scenario and attribution protocols

Two instruments published before use: a scenario-based protocol for pluralistic moral reasoning in model outputs, and a human-attribution rating protocol. Both are labelled unvalidated, because they are.

Two records
Governance & synthesis

Outputs are not minds, but they still matter

A synthesis arguing that downstream consequences come from output cues, human attribution and institutional incentives jointly — with five conditional forecasts stated so they can later be checked.

Perspective

About

How I work.

I work on evaluation, measurement and reliability in language-model behaviour. The work runs from the question to the released artifact: designing the study, building the harness, running it, auditing it, finding out what broke, and writing it up in a form somebody else can check.

What holds it together is a preference for methods that can return a negative. The programme on this site began with an observation I was not neutral about, which is a bad place to start reasoning from and a good place to start building instruments from. Most of the engineering since has been about making it possible for the evidence to contradict me.

In practice that means the answer condition is written down before the data exists, the checks are implemented as code that fails closed rather than as a checklist, and a result that comes back indeterminate gets published as indeterminate.

Where this maps

AI evaluation · model behaviour · reliability and assurance · safety evaluations · evaluation infrastructure · research software engineering · interpretability research engineering · reproducible ML experimentation.

How this work was produced: the division of labour

This research was produced with substantial help from modern AI coding and research systems, and every deposited paper discloses that on its first page. Neither “an AI did it” nor “one person did all of it” would be an accurate description, so here is the actual split.

What I decided

  • Which question was worth asking, and how to decompose it into testable parts
  • How each concept was operationalised, and what the resulting number would and would not license
  • The acceptance criteria: thresholds, routes, and what happens when a threshold is missed
  • When a result was insufficient and the method had to be redesigned rather than the bar lowered
  • Which conceptual failures mattered — including the ones only visible across analyses
  • What required independent verification, and what the final claim boundaries would be

What AI systems substantially assisted with

  • Implementation: harness code, capture pipelines, analysis and statistics modules
  • Critique: adversarial review passes over methodology, code and manuscripts
  • Analysis: recomputation, reanalysis and cross-checking of reported figures
  • Review: drafting, editing and consistency checking across a large corpus

The part worth paying attention to is neither side of that table but the structure between them: nothing on either side is trusted by default. Model output is checked against the test suite, the manifests, the recomputed statistics and the source artifacts; my own preferred interpretation is checked against a threshold I fixed before I could see the result. On more than one occasion an automated critique pass found a real defect in my method, and the method changed.

That is the workflow I would bring to a team: heavily leveraged AI tooling, with the verification externalised into artifacts that can contradict both the tool and the person using it.

  • Fix the criterion firstThresholds and routes are frozen before collection, so a disappointing result cannot be renegotiated into a publishable one.
  • Build the check into the codeA convention gets forgotten. A process that exits non-zero does not.
  • Separate observation from interpretationWhat the output did and what it might mean are different claims, and the second one is usually the expensive one.
  • Keep the negative resultsA null that survived its controls is information. Two of the three case studies here are negative, and they are the ones I would want read.
  • Make it recomputableManifests, pinned revisions, deterministic seeds. If the bytes do not verify, the claim does not ship.

What I can contribute

What this would look like on a team.

Six things I have done end to end, stated as what they would be useful for. Each one names the work it comes from.

Evaluation design

Turn an ambiguous question about model behaviour into a testable procedure with a pass mark fixed before collection — including what happens when the pass mark is missed.

From: a 64-block pre-registered behavioural assay and its reliability gate.

Research engineering

Build and maintain experiment harnesses, capture and analysis pipelines, and workflows where a stage refuses to run until its predecessor verifies.

From: the evaluation harness and the public reproducibility package.

Failure analysis

Investigate an unexpected result and localise it, without upgrading the observation into a conclusion on the way — including testing whether it survives deduplication and dependence.

From: the stratified reliability diagnosis and the dependence-aware reanalysis.

Interpretability evaluation

Capture activations, screen a wide feature space, and test whether an internal signal survives the alternative explanations that would account for it just as well.

From: the five-gate cascade over 65,536 layer-feature coordinates.

Research reliability

Build the systems that preserve provenance, validate artifacts and make a result independently inspectable — manifests, write-once publication, pinned environments, deterministic seeds.

From: the release verification system and its independently recomputed numbers.

High-leverage AI tooling

Direct AI coding and research systems on substantial work while keeping the verification external — in tests, manifests, recomputation and pre-committed criteria.

From: the whole programme, disclosed on every deposited paper.

Contact

Open to evaluation and research-engineering work.

If the work involves making a measurement trustworthy before making it impressive, I would like to hear about it. Evaluation, model behaviour, reliability and assurance, safety evaluations, evaluation infrastructure, research software engineering and interpretability research engineering.