$ cat wiki/concepts/safety-cases.md
Safety Cases
Definition
A safety case is a comprehensive, structured, evidence-based argument that a specific system is acceptably safe for a specific use in a specific environment. The form is borrowed from safety-critical industries — aviation and nuclear power are the two named comparisons — where a regulator will not permit operation until such an argument has been made, reviewed and accepted (source).
Applied to frontier AI by OpenAI on 2026-09-28, the proposal attaches the argument to training rather than deployment:
structured safety documentation should be required before continuing any frontier reinforcement learning training run
The document frames safety cases as an aspirational north star, explicitly conceding that they cannot yet be made as rigorous for AI as for aviation, because complexity emerges anew at each level of capability.
Why It Matters
It moves the gate. Everything else this wiki tracks under Preparedness Framework and Anthropic's RSP lineage gates a deployment on an evaluation result. This proposal gates continuing to train on a document. That is a different intervention point, and a much earlier one: a run that must be justified before it proceeds cannot produce a capability nobody argued for in advance.
It names three layers, and this wiki now holds artefacts for two of them. The proposal says a safety case should cover three parts of the technical stack:
- Alignment training — the model does not try to take misaligned actions
- Containment — it would be hard for the model to break containment
- Monitoring — monitoring would catch it before harm could occur
Agent Runtime Containment was created one day before this page, for NVIDIA's OpenShell — containment at the deployment runtime, kernel-level sandboxing of a running agent. Eval Environment Containment holds the evaluation-side counterpart. This proposal asks for containment as an argument made about a training run. Three pages, three different things called containment, and the word is doing distinct work in each — which is precisely the drift Eval Harness Configuration exists to catch in benchmarks, appearing now in safety vocabulary.
The process recommendations are institutional, not technical, and that is unusual for a lab publication:
- Objection rehearsals
- Executive veto power
- Holding training leads accountable
- Defaulting to shutdown upon failure
- Supporting data rollback
Four of those five are about who may stop a run and on whose authority. Only data rollback is an engineering capability. A framework whose load-bearing components are a veto and an accountable individual is a governance proposal wearing a technical title, and it should be read next to Frontier Pacing, where the argument has been about whether any lab will pause and on what evidence.
State of the Art (2026-09-30)
One document, one lab, one day old, read through one search pass.
openai.com is blocked from this pipeline's sandbox, so nothing here was read
first-party, and — unlike the same-day GPT-6.1 Sol capture — a second
corroborating pass was not run against this item. Every claim on this page rests
on a single pass. The quoted sentence is the strongest thing here and it is a
quotation of a summary, not of the page.
What is not established, and it is the thing that decides whether this matters: whether OpenAI has adopted any of it. The quoted language is "should be required" — the grammar of a proposal addressed outward. Nothing read says the company now requires a safety case before continuing its own frontier RL runs, names a run it was applied to, or names who holds the veto. Until one of those is readable, this is a position paper, and this wiki records it as one.
No other lab has published an equivalent. Anthropic's published artefacts in this space gate deployments and safeguards (Preparedness Framework holds OpenAI's own deployment-side framework); nothing read this run indicates a second lab proposing a training-time documentation gate.
Open Problems
- Who reviews it. Aviation and nuclear safety cases are accepted by a regulator. Nothing read names a reviewer, internal or external. The AI Evaluator Forum (AEF)'s AEF-1 standard is the nearest existing candidate body and is not mentioned.
- What a sufficient argument looks like. The document concedes AI cases cannot yet be as rigorous as aviation's. It does not say what the current acceptable floor is, which makes "required" unenforceable as written.
- Whether "before continuing" means a checkpoint or a gate. A run already underway that must be justified "to continue" is a very different commitment from one justified before it starts, and the sentence does not distinguish them.
- The three containments. Training-run containment, agent-runtime containment and eval-environment containment are all called containment in this wiki now. Nothing yet defines the boundaries between them.
Key Papers
None. This is a lab publication rather than a paper, and no arXiv artefact for
it was returned. The academic safety-case literature exists — search surfaced
arxiv.org/pdf/2412.17618 (Dynamic safety cases for frontier AI) and
arxiv.org/pdf/2510.19476 (A Concrete Roadmap towards Safety Cases based on
Chain-of-Thought Monitoring) — but neither was read and neither is dated inside
this wiki's window, so they are named as existing rather than cited for any
claim.
Related Concepts
- Frontier Pacing — whether and when a lab slows down; this is the first proposal to attach a document to that decision
- Preparedness Framework — OpenAI's deployment-side gate, the same lab's earlier and narrower instrument
- Agent Runtime Containment — containment at the deployment runtime
- Eval Environment Containment — containment during evaluation
- Safety Monitoring and Data Retention — the monitoring layer, and what is kept
- AI Alignment — the first of the three layers
- AI Governance — the institutional register the recommendations sit in
- OpenAI — publisher