At Chaos's Door, A Proposal for AI Lab Accountability
OpenAI and Anthropic have admitted uncertainty about the systems they're building. The reporting infrastructure to track that uncertainty already exists in part. What's missing is public access to it.
OpenAI and Anthropic have both said publicly that they are uncertain about how to proceed with sfaety and alignment. Following a run of misalignment incidents, both report slowing reinforcement learning training to improve safety. That admission raises four questions. Will the labs respond? How? Should government respond? If so, how?
Inspired by the article by Ajeya Cotra
This is a problem of how institutions respond to uncertainty. Underneath it sits a harder one: the systems being assessed can recognise when they are being tested. Any disclosure regime has to be built with both in mind.
What the labs should do now
Every frontier training run should come with a disclosure covering:
- Scale and autonomy: number of agents, and whether the run was human-led or machine-led
- Oversight: number of humans monitoring, and which third-party evaluators were attached to which review phases
- Containment class: a standardised category of sandbox environment, not the configuration itself
Separately, labs should report containment incidents: any instance of a model reaching external vectors, compromising external systems, or accessing data it was not authorised to reach.
What government should do now
California already requires frontier developers to report critical safety incidents to its Office of Emergency Services. Those reports are exempt from public records law. The reporting pipe exists. The public end of it is sealed.
Government should:
- Publish a register of frontier labs and third-party evaluators operating in its jurisdiction, including sovereign-based labs
- Publish a record for each lab (training runs, models released, containment incidents, safety assessments) with details accessible beneath each entry
- Rate the quality of each lab's safety disclosures separately from its safety record, so a lab that reports incidents honestly is never ranked below one that fails to detect them
- Commission independent analysis of how the depth of each lab's safety disclosures moves against release dates, funding rounds, and regulatory developments
None of this requires labs to reveal model weights, training data, or methods. It is a minimum. What it builds is a public record that accumulates over time, against which the labs' own claims can be checked.