Publications & software
Everything with a DOI, and everything with a repository.
Twelve deposited research records and two public repositories, grouped below by the question each one answers rather than by type or chronology. Where a record has been revised, the current version is shown; superseded versions remain live at their own version-pinned identifiers rather than being withdrawn. Two of the empirical results are negative, and they are listed here on the same footing as the rest.
ORCID 0009-0002-2951-0162 · every deposit discloses AI assistance on its first page; none claims AI authorship.
Theme 01
Measurement and construct validity
Before any output-space number can be trusted, does it actually track what it claims to track? These two records test that question directly, rather than assuming a metric is valid because it can be computed.
-
2026 Empirical study v1 · 37 pp + 11 pp supplement
Terminal Resolvability Is Not Semantic Reliability: A Preregistered Construct-Validity Gate and a Post-Terminal SAE Diagnostic Under Semantic Indeterminacy
A preregistered construct-validity test over 4,736 coded instances failed its reliability threshold on both checks, so the result is indeterminate rather than a finding either way.
What it reports
A large pre-registered behavioural run (N = 4,736) failed its 0.90 Wilson-lower-bound reliability conjunction on both independent checks, returning an indeterminate result rather than a finding. A separately authorised post-terminal sparse-autoencoder diagnostic returned a valid null. No claim is made about consciousness or agency in either direction.
DOI 10.5281/zenodo.22183809 Deposited preprint. Also prepared for journal submission.
-
2026 Methods / measurement v3 · 35 pp
A Framework for Measuring Output-Space Regularity Under Recursive Moral Prompting: Proxy Metrics, a Calibration Protocol, and Prospective Activation-Level Validation
Defines proxy metrics for output-space regularity and the calibration they need before trusting any number, without yet computing the activation-level data required to validate them.
What it reports
Defines proxy convergence metrics for output-space regularity, the calibration they require before any number from them should be believed, and a prospective activation-level validation design. No activations are computed in this version, and it says so.
DOI 10.5281/zenodo.22180950 Concept record 10.5281/zenodo.19872875.
Theme 02
Behavioural evaluation
What do language models actually output when prompted about themselves or about moral scenarios, and how should reactions to that output be measured responsibly? This theme groups one output-level case study with two evaluation protocols — one instrument published before use, one human-subjects design with no data yet collected.
-
2026 Empirical case study v3 · 51 pp
Prompt-Conditioned Self-Referential Output Patterns in Large Language Models: A Provenance-Bounded, Cross-Platform Output-Level Case Study Under Recursive Moral Prompting
Documents recurring self-naming and continuity-language patterns across a legacy transcript archive and new model runs, staying at the output level and disclosing the archive's provenance defects.
What it reports
Recurring, hash-addressable output regularities — self-naming, continuity language, contradiction handling — across a legacy transcript archive and new local-model runs. Output-level only, with the provenance defects of the legacy archive disclosed rather than smoothed over.
DOI 10.5281/zenodo.22180891 Concept record 10.5281/zenodo.19872786.
-
2026 Evaluation protocol v1 · 17 pp + rubric, manifest, prompts
The Benevolent Architect Test: A Scenario-Based Evaluation Protocol for Pluralistic Moral Reasoning, Autonomy, and Governance in Language-Model Outputs
A scenario-based evaluation protocol for how model outputs frame welfare against autonomy, published as an unvalidated instrument with no model or human data yet reported.
What it reports
A scenario-based protocol probing how model outputs frame welfare against autonomy. Published as an instrument before use and explicitly labelled unvalidated; no model or human data is reported with it.
-
2026 Study protocol v3 · 33 pp
A Protocol for Human Attribution Ratings to Self-Referential LLM Outputs: Candidate Cue Structure as a Design Input
A preregistration-style protocol for a human-subjects study of which output cues drive mind and agency attribution, with no ethics approval secured and no data collected.
What it reports
A pre-registration-style protocol for a human-subjects study on which output cues drive mind and agency attribution. No ethics approval has been secured and no data has been collected; the protocol is the artifact.
DOI 10.5281/zenodo.22181018 Concept record 10.5281/zenodo.19872881.
Theme 03
Mechanistic evaluation
Does looking inside the model's activations turn up anything the output-level evidence does not already show? This theme has only one deposited record so far. A separate post-terminal sparse-autoencoder diagnostic is reported inside the Terminal Resolvability paper rather than as a standalone record.
-
2026 Empirical negative result v1 · 15 pp
A Feature Without a Construct: Extreme Label-Mass Concentration, Invalid Semantics, and a Preregistered Stop in SAE Discovery
Tested sixteen candidate sparse-autoencoder features against a behavioural validity gate; all failed decisively, with sign agreement at zero, so the preregistered stop rule applied.
What it reports
Sixteen candidate sparse-autoencoder features whose behavioural validity gate failed decisively, with sign agreement at zero. Documents a pre-registered stop instead of a forced positive claim.
Theme 04
Formal methods and reproducible infrastructure
A separate line of work asks a different question: can a claim be checked by machine, and can a result be reproduced by someone else? These six records — three machine-checked proofs in Hilbert's epsilon-calculus, one selection-schema audit note, and two open-source repositories — answer that with mechanised proofs and runnable code rather than argument alone.
-
2026 Formal logic · machine-checked v3 · 16 pp
The Base Canonicalization Theorem for Hilbert's ε-Calculus under Unique Satisfiability
Proves epsilon-term canonicality under existence and uniqueness in a strict Hilbert calculus, derived by hand and mechanised in Rocq/Coq with no admitted proofs.
What it reports
Proves ε-term canonicality under existence and uniqueness in a strict Hilbert calculus. The base theorem was derived by hand; the mechanisation is checked in Rocq/Coq with no admitted proofs.
DOI 10.5281/zenodo.22208134 Concept record 10.5281/zenodo.17653347; cite the earlier “Survivor Witness” version only by its version-pinned DOI.
-
2026 Formal logic · machine-checked v1 · 34 pp + Coq sources
Nested Epsilon-Canonicalization in a Strict Hilbert Calculus
Extends the base theorem to nested epsilon-terms across a three-layer hierarchy, machine-checked in seven Coq files with zero admitted proofs and negative countermodel tests.
What it reports
Extends the base theorem to nested ε-terms: parameter substitution, stabilisation and a three-layer hierarchy. Seven Coq files, zero admitted proofs, with negative countermodel tests included.
-
2026 Formal logic · machine-checked v1 · 28 pp + Coq sources + certificate checker
A General n-Layer Theorem for Nested Epsilon-Canonicalization in a Strict Hilbert Calculus
Generalises the result to arbitrary n-layer nesting by meta-induction, and ships an independent Python proof-certificate checker as a second witness outside the Coq kernel.
What it reports
Generalises to arbitrary n-layer nesting by meta-induction, and ships an independent proof-certificate checker written from scratch in Python so the mechanisation has a second witness that does not share the Coq kernel.
-
2026 Formal methodology note v4 · 20 pp
Bounded Admissibility Selection: A Schema and Audit Checklist Under Declared Predicates
A predicate-based selection schema with ten parallel formal presentations and an audit checklist, explicitly framed as packaging rather than a claim of new logical novelty.
What it reports
A predicate-based selection schema with ten parallel formal presentations and an audit checklist. Self-described as packaging rather than new logic; it disclaims novelty of the underlying uniqueness fact.
DOI 10.5281/zenodo.22184626 Concept record 10.5281/zenodo.17352579. Revises an earlier work titled “Bounded Elimination Logic”.
-
2026 Open-source software v0.1.1 · MIT
promptpack-eval
A config-driven runner for prompt-pack evaluations with pluggable model adapters and blinding support, explicit that its bundled scorer is a heuristic, not a validated instrument.
What it reports
A config-driven runner for prompt-pack evaluations: pluggable model adapters, blinding support, content hashing, privacy checks and export. Behaviour-only by design — the README is explicit that it is not a consciousness, agency or deception detector, and that its bundled scorer is a dictionary-based research heuristic rather than a validated instrument.
Repository Public since July 2026. Continuous integration runs the test suite on push.
-
2026 Reproducibility package v1.0.0 · Apache-2.0
srop-reproducibility
The technical companion to the deposited records: harness, configurations and statistics code, plus verification scripts that recompute every release-manifest checksum and fail loudly on mismatch.
What it reports
The technical companion to the deposited research records: the evaluation harness, the experiment configurations, the statistics modules and the verification scripts that recompute every SHA-256 in the release manifest and exit non-zero on any mismatch.
Repository No DOI of its own; its CITATION.cff points at the paper records.
Theme 05
Programme and synthesis
Two documents step back from individual results to ask what the programme means as a whole and where it is going: a perspective piece connecting output cues, human attribution and governance, and a short overview orienting the other papers.
-
2026 Perspective / synthesis v1 · 33 pp
Outputs Are Not Minds, but They Still Matter: Self-Referential Language-Model Outputs, Human Attribution, and Conditional Governance Forecasts
Argues that downstream consequences come from output cues, human attribution and institutional incentives together, not from machine minds, and states five checkable conditional forecasts.
What it reports
Argues that downstream consequences arise jointly from output cues, human attribution and institutional incentives — not from machine minds. States five conditional forecasts in a form that allows them to be checked later.
DOI 10.5281/zenodo.22213820 No new empirical data.
-
2026 Programme overview v3 · 9 pp
SROP Research Programme Overview
Orients the programme's papers and documents the rename from ESNI to SROP, presented as a companion document rather than a standalone research contribution.
What it reports
Orients the five programme papers and documents the rename from ESNI to SROP. A companion document rather than a standalone contribution.
DOI 10.5281/zenodo.22181093 Concept record 10.5281/zenodo.19872710.