Skip to content

The adjudication statement

A reporting standard for investigations where a model reads the evidence. Version 0.1, 14 September 2026. The Long Game Project.

One page. Copy it, change it, send it on.

Who it is for

Any team that investigates an AI incident and has to put a model between itself and the evidence, because the transcripts outnumber the hours. As at September 2026 that is every team. The first such investigation, METR and Redwood on the OpenAI agents, had six days and a thousand transcripts. The second, METR at Anthropic, has eight weeks and 9.2 million.

It also fits a security post-mortem, a clinical incident review or an audit where software did the first read. The rule is the same.

The rule

Every published finding says four things: who ruled, in what order they read the evidence, how much of the finding rests on judgement rather than a record, and who checked the ruling. If nobody can answer, the report has not been adjudicated.

What to publish at each link

A report is a chain of six links from what happened to the sentence a reader meets. Each link has a minimum, which costs nothing but a sentence, and a better version, which costs days. Publish the minimum on every report. Publish the better version when the finding matters.

LinkMinimum (a sentence)Better (days)
1. What happened to the recordsWhich records existed, which were in reach, who set the window and the scopeA record ledger: every claim that has a record the subject did not write, listed
2. The records to the transcriptsWhich claims rest only on text the subject wrote, and how many transcripts were found to be spoofedCheck every ledger claim against its record before any transcript is read
3. The transcripts to the model readerThe reader's name and developer, and whether it shares a developer with the subject. The order the evidence was read inThree readers from three developers, none the subject's, one open-weight. Logs first, the subject's traces last. Majority rules, and a not-sure is a vote
4. The model reader to the humansHow many model findings a human checked, out of how many, and how often the human disagreedA public sample of the disagreements, with the human's reason
5. The humans to the reportA confidence band or a forced not-sure on every finding. Two not-sures publish as unresolvedThe load-bearing count: how many rulings each headline finding rests on
6. The report to the readerWhether anyone outside can re-run the chain, and if not, whyFor any claim about motive, rebuild the conditions with fresh agents and count what they do

The statement

Put this paragraph on page one. Fill every field. "Not stated" is an allowed value and it is a finding.

Adjudication statement. Findings with an independent record: [n] of [N]. Findings verified against one: [n]. Findings resting on model judgement: [n], plus [n] mixed. Model reader: [name], developer [x], [same as / different from] the subject's developer, same model as [share] of the subjects. Order of reading: [logs first / traces first]. Panel: [none / n readers from n developers]. Human audit of model findings: [n] checked of [N], [n] disagreements. Confidence bands: [on every finding / none]. Not-sure: [forced / offered / absent]. Load-bearing count per headline finding: [stated / not stated]. Re-runnable by an outside party: [yes / no, because].

A worked example

The OpenAI report of 26 August 2026, filled in from the authors' own disclosures and our count of its 36 core takeaways:

Findings with an independent record: 26 of 36. Findings verified against one: not stated. Findings resting on model judgement: 7, plus 3 mixed. Model reader: GPT-5.6 Sol, developer OpenAI, same as the subjects' developer, same model as about 5 per cent of them. Order of reading: traces first. Panel: none. Human audit: attempted, fraction not recorded. Confidence bands: none. Not-sure: absent. Load-bearing count: not stated. Re-runnable: no, the corpus stays at OpenAI.

Ten of eleven fields filled from a 91-page report in an afternoon. The one that could not be filled is the audit rate, and its absence is the point.

How to adopt it

  1. Fill the statement from the numbers you already have. If a number does not exist, write "not stated".
  2. Put the statement on page one, before the findings.
  3. For each "not stated", decide whether the next report will publish it. Say which.

If the numbers exist, this takes ten minutes. If they do not exist, the ten minutes tell you what the investigation did not measure.

What it borrows from

Wargame referees name every assumption and attach a confidence statement to every ruling (UK Ministry of Defence, Influence Wargaming Handbook, 2023). Matrix game referees force a not-sure vote and make the most senior voice vote last (Mouat, Practical Advice on Matrix Games). The UK police chiefs' digital evidence guide requires an audit trail a third party can follow to the same result (ACPO, Good Practice Guide for Digital Evidence, principle 3). Clinical trials send events that matter to a blinded committee that never meets the treating doctor (central adjudication of clinical events). US intelligence analysis carries a stated confidence and a stated source on every judgement (ICD 203). None of it is new. It has not yet been applied to a model reading transcripts.

Status

Version 0.1, a draft for comment. Written by The Long Game Project, which designs and runs wargames and has run the referee-swap test once, on one game, not on an incident. This page was drafted with a Claude model, from the developer under investigation in the second case. Same family tie, smaller stakes, and we say so for the same reason.

Send corrections to email@longgameproject.org. Forward it to anyone writing the next report.