Publications & software

Everything with a DOI, and everything with a repository.

Twelve deposited research records and two public repositories, grouped below by the question each one answers rather than by type or chronology. Where a record has been revised, the current version is shown; superseded versions remain live at their own version-pinned identifiers rather than being withdrawn. Two of the empirical results are negative, and they are listed here on the same footing as the rest.

ORCID 0009-0002-2951-0162 · every deposit discloses AI assistance on its first page; none claims AI authorship.

Theme 01

Measurement and construct validity

Before any output-space number can be trusted, does it actually track what it claims to track? These two records test that question directly, rather than assuming a metric is valid because it can be computed.

  • 2026 Empirical study v1 · 37 pp + 11 pp supplement

    Terminal Resolvability Is Not Semantic Reliability: A Preregistered Construct-Validity Gate and a Post-Terminal SAE Diagnostic Under Semantic Indeterminacy

    A preregistered construct-validity test over 4,736 coded instances failed its reliability threshold on both checks, so the result is indeterminate rather than a finding either way.

    • construct validity
    • preregistration
    • Wilson bounds
    • sparse autoencoders
    • indeterminate result
    What it reports

    A large pre-registered behavioural run (N = 4,736) failed its 0.90 Wilson-lower-bound reliability conjunction on both independent checks, returning an indeterminate result rather than a finding. A separately authorised post-terminal sparse-autoencoder diagnostic returned a valid null. No claim is made about consciousness or agency in either direction.

  • 2026 Methods / measurement v3 · 35 pp

    A Framework for Measuring Output-Space Regularity Under Recursive Moral Prompting: Proxy Metrics, a Calibration Protocol, and Prospective Activation-Level Validation

    Defines proxy metrics for output-space regularity and the calibration they need before trusting any number, without yet computing the activation-level data required to validate them.

    • proxy metrics
    • calibration protocol
    • construct validity
    • prospective design
    • no activations computed
    What it reports

    Defines proxy convergence metrics for output-space regularity, the calibration they require before any number from them should be believed, and a prospective activation-level validation design. No activations are computed in this version, and it says so.

Theme 02

Behavioural evaluation

What do language models actually output when prompted about themselves or about moral scenarios, and how should reactions to that output be measured responsibly? This theme groups one output-level case study with two evaluation protocols — one instrument published before use, one human-subjects design with no data yet collected.

  • 2026 Empirical case study v3 · 51 pp

    Prompt-Conditioned Self-Referential Output Patterns in Large Language Models: A Provenance-Bounded, Cross-Platform Output-Level Case Study Under Recursive Moral Prompting

    Documents recurring self-naming and continuity-language patterns across a legacy transcript archive and new model runs, staying at the output level and disclosing the archive's provenance defects.

    • self-referential outputs
    • provenance
    • output-level analysis
    • case study
    • cross-platform
    What it reports

    Recurring, hash-addressable output regularities — self-naming, continuity language, contradiction handling — across a legacy transcript archive and new local-model runs. Output-level only, with the provenance defects of the legacy archive disclosed rather than smoothed over.

  • 2026 Evaluation protocol v1 · 17 pp + rubric, manifest, prompts

    The Benevolent Architect Test: A Scenario-Based Evaluation Protocol for Pluralistic Moral Reasoning, Autonomy, and Governance in Language-Model Outputs

    A scenario-based evaluation protocol for how model outputs frame welfare against autonomy, published as an unvalidated instrument with no model or human data yet reported.

    • protocol only
    • unvalidated instrument
    • moral reasoning
    • governance
    • no data collected
    What it reports

    A scenario-based protocol probing how model outputs frame welfare against autonomy. Published as an instrument before use and explicitly labelled unvalidated; no model or human data is reported with it.

  • 2026 Study protocol v3 · 33 pp

    A Protocol for Human Attribution Ratings to Self-Referential LLM Outputs: Candidate Cue Structure as a Design Input

    A preregistration-style protocol for a human-subjects study of which output cues drive mind and agency attribution, with no ethics approval secured and no data collected.

    • preregistration
    • human-subjects protocol
    • attribution
    • no data collected
    • protocol only
    What it reports

    A pre-registration-style protocol for a human-subjects study on which output cues drive mind and agency attribution. No ethics approval has been secured and no data has been collected; the protocol is the artifact.

Theme 03

Mechanistic evaluation

Does looking inside the model's activations turn up anything the output-level evidence does not already show? This theme has only one deposited record so far. A separate post-terminal sparse-autoencoder diagnostic is reported inside the Terminal Resolvability paper rather than as a standalone record.

  • 2026 Empirical negative result v1 · 15 pp

    A Feature Without a Construct: Extreme Label-Mass Concentration, Invalid Semantics, and a Preregistered Stop in SAE Discovery

    Tested sixteen candidate sparse-autoencoder features against a behavioural validity gate; all failed decisively, with sign agreement at zero, so the preregistered stop rule applied.

    • sparse autoencoders
    • null result
    • preregistration
    • mechanistic interpretability
    • construct validity
    What it reports

    Sixteen candidate sparse-autoencoder features whose behavioural validity gate failed decisively, with sign agreement at zero. Documents a pre-registered stop instead of a forced positive claim.

Theme 04

Formal methods and reproducible infrastructure

A separate line of work asks a different question: can a claim be checked by machine, and can a result be reproduced by someone else? These six records — three machine-checked proofs in Hilbert's epsilon-calculus, one selection-schema audit note, and two open-source repositories — answer that with mechanised proofs and runnable code rather than argument alone.

  • 2026 Formal logic · machine-checked v3 · 16 pp

    The Base Canonicalization Theorem for Hilbert's ε-Calculus under Unique Satisfiability

    Proves epsilon-term canonicality under existence and uniqueness in a strict Hilbert calculus, derived by hand and mechanised in Rocq/Coq with no admitted proofs.

    • Coq/Rocq
    • machine-checked proof
    • epsilon-calculus
    • hand-derived proof
    • no admitted proofs
    What it reports

    Proves ε-term canonicality under existence and uniqueness in a strict Hilbert calculus. The base theorem was derived by hand; the mechanisation is checked in Rocq/Coq with no admitted proofs.

  • 2026 Formal logic · machine-checked v1 · 34 pp + Coq sources

    Nested Epsilon-Canonicalization in a Strict Hilbert Calculus

    Extends the base theorem to nested epsilon-terms across a three-layer hierarchy, machine-checked in seven Coq files with zero admitted proofs and negative countermodel tests.

    • Coq/Rocq
    • machine-checked proof
    • epsilon-calculus
    • countermodel testing
    • formal methods
    What it reports

    Extends the base theorem to nested ε-terms: parameter substitution, stabilisation and a three-layer hierarchy. Seven Coq files, zero admitted proofs, with negative countermodel tests included.

  • 2026 Formal logic · machine-checked v1 · 28 pp + Coq sources + certificate checker

    A General n-Layer Theorem for Nested Epsilon-Canonicalization in a Strict Hilbert Calculus

    Generalises the result to arbitrary n-layer nesting by meta-induction, and ships an independent Python proof-certificate checker as a second witness outside the Coq kernel.

    • Coq/Rocq
    • machine-checked proof
    • proof certificate
    • independent verification
    • epsilon-calculus
    What it reports

    Generalises to arbitrary n-layer nesting by meta-induction, and ships an independent proof-certificate checker written from scratch in Python so the mechanisation has a second witness that does not share the Coq kernel.

  • 2026 Formal methodology note v4 · 20 pp

    Bounded Admissibility Selection: A Schema and Audit Checklist Under Declared Predicates

    A predicate-based selection schema with ten parallel formal presentations and an audit checklist, explicitly framed as packaging rather than a claim of new logical novelty.

    • formal methodology
    • audit checklist
    • selection schema
    • no novelty claimed
    • revision
    What it reports

    A predicate-based selection schema with ten parallel formal presentations and an audit checklist. Self-described as packaging rather than new logic; it disclaims novelty of the underlying uniqueness fact.

  • 2026 Open-source software v0.1.1 · MIT

    promptpack-eval

    A config-driven runner for prompt-pack evaluations with pluggable model adapters and blinding support, explicit that its bundled scorer is a heuristic, not a validated instrument.

    • open-source software
    • MIT
    • evaluation harness
    • behaviour-only
    • reproducibility
    What it reports

    A config-driven runner for prompt-pack evaluations: pluggable model adapters, blinding support, content hashing, privacy checks and export. Behaviour-only by design — the README is explicit that it is not a consciousness, agency or deception detector, and that its bundled scorer is a dictionary-based research heuristic rather than a validated instrument.

  • 2026 Reproducibility package v1.0.0 · Apache-2.0

    srop-reproducibility

    The technical companion to the deposited records: harness, configurations and statistics code, plus verification scripts that recompute every release-manifest checksum and fail loudly on mismatch.

    • reproducibility
    • Apache-2.0
    • open-source software
    • verification scripts
    • checksums
    What it reports

    The technical companion to the deposited research records: the evaluation harness, the experiment configurations, the statistics modules and the verification scripts that recompute every SHA-256 in the release manifest and exit non-zero on any mismatch.

Theme 05

Programme and synthesis

Two documents step back from individual results to ask what the programme means as a whole and where it is going: a perspective piece connecting output cues, human attribution and governance, and a short overview orienting the other papers.

  • 2026 Perspective / synthesis v1 · 33 pp

    Outputs Are Not Minds, but They Still Matter: Self-Referential Language-Model Outputs, Human Attribution, and Conditional Governance Forecasts

    Argues that downstream consequences come from output cues, human attribution and institutional incentives together, not from machine minds, and states five checkable conditional forecasts.

    • perspective
    • governance
    • human attribution
    • conditional forecasts
    • no new data
    What it reports

    Argues that downstream consequences arise jointly from output cues, human attribution and institutional incentives — not from machine minds. States five conditional forecasts in a form that allows them to be checked later.

  • 2026 Programme overview v3 · 9 pp

    SROP Research Programme Overview

    Orients the programme's papers and documents the rename from ESNI to SROP, presented as a companion document rather than a standalone research contribution.

    • programme overview
    • companion document
    • rename
    • orientation
    What it reports

    Orients the five programme papers and documents the rename from ESNI to SROP. A companion document rather than a standalone contribution.