$ cat wiki/concepts/safety-monitoring-retention.md
Safety Monitoring and Data Retention
Definition
Whether a frontier lab must hold customer prompts and outputs in order to detect misuse of its most capable models — and, if it must, for how long, under what access control, and with what right of refusal for the customer.
The question only became live when detection moved from per-request to cross-request. A classifier that scores a single prompt needs nothing after the request completes. Detecting a pattern across related interactions — the abuse shape both labs now say matters most at the frontier — requires comparing a request against others, which requires having kept them (source).
Zero Data Retention (ZDR) is the enterprise commitment that the provider keeps nothing after a request is processed. It predates this problem and was, until 2026, an ordinary procurement checkbox. The frontier-model safety case is the first thing to contest it.
Why It Matters
Two frontier labs have now taken opposite architectural positions on the same problem, within ten weeks of each other, and each says its position is required for safety:
| Anthropic | OpenAI | |
|---|---|---|
| Position | retention is necessary for safety — withdrawn 2026-09-01 | retention is not necessary for safety |
| Mechanism | Enterprise Frontier Safeguards — automated monitoring over logs held in the customer's own cloud; superseding the 30-day retention of all Mythos-class traffic | Private Safety Processing — a narrow signal, no content exposure |
| Effect on a customer's existing ZDR agreement | voided 2026-06-09 with no opt-out; zero data retention restored under EFS | preserved; ZDR granted on prior approval |
| Scope | all traffic, first- and third-party surfaces | eligible enterprise/API endpoints only |
| Status | announced 2026-09-01, phased rollout from fall 2026; the 06-09 policy was in force for ~12 weeks | previewed with early customers; rollout and white paper due September 2026 |
| Consumer plans | unaffected — already retained | not covered |
| Sources: | ||
| (Anthropic), | ||
| (OpenAI), | ||
| (Anthropic change, reported). |
The asymmetry in what each has shipped moved within twenty-four hours of this page being written, and the direction is the news. On 2026-08-20 Bloomberg reported that Anthropic plans to let enterprise customers hold the retained data on their own cloud infrastructure, with the 30-day requirement itself unchanged (source). Anthropic's side is therefore also a plan now — and it converges on the mechanism OpenAI previewed the day before, customer-controlled infrastructure.
What survives the convergence is the premise, and it is the thing worth tracking. Both labs still assert that frontier risk shows up across interactions rather than in a single prompt–response pair; both now propose that the customer can hold the content. The disagreement narrows from "must content be retained" to "who holds it, and what does the lab receive" — and neither has said what the lab receives. That is a smaller question than the one this page was created for, and a sharper one.
And the disagreement is technical, not merely commercial. Both labs accept the same premise — that frontier-model risk shows up across interactions rather than in a single prompt–response pair — and disagree about whether that premise forces content retention. Only one of them can be right, and the answer determines whether "zero retention" survives as a category for frontier models at all.
State of the Art (as of 2026-09-08)
The question this page has treated as a policy argument became a specific accusation about a specific person's sessions, and the lab's answer has a carve-out in it. In the credit dispute over the 2026-09-08 Navier–Stokes result — treated as mathematics on AI for Mathematics, and taken here only for the data question — Tristan Buckmaster said publicly that he wonders whether the private mathematical drafts he had been feeding into OpenAI's tools were accessible to OpenAI (3 passes). OpenAI's reply, quoted in two independent passes, is two sentences and they do different work (source):
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
The first sentence denies access; the second declines to deny influence, and the boundary between them is exactly the one this page exists to track. Everything on this page above concerns retention for misuse detection — what is kept, by whom, for how long, so that a classifier can compare a request against others. This is the same substrate reached by a different route: training. A retention regime designed around who may read a log says nothing about whether a de-identified derivative of it has already moved into weights, and "we cannot rule out" is the strongest statement a lab can make when the two pipelines are separate and only one of them has an access log.
Three things are worth stating plainly about the evidence. Nothing read contains a log, an audit or a third-party examination — OpenAI's denial is a statement and Buckmaster's concern is a question, and neither is corroborated. The applicable retention terms for the product he describes using were not quoted in anything read, so this wiki cannot say what was promised, let alone what happened. And no Anthropic statement appears anywhere in the coverage, though Alpöge is identified as an Anthropic employee in 3 passes — the lab whose enterprise offer is the zero-retention product described below has said nothing about a case that turns on a competitor's retention.
What it changes here is the frame, not a fact. Open Problem 4 asks whether a retention duration still applies to customer-held logs. This adds a question that duration cannot answer: whether de-identified derivatives outlive the retention window by design, and what a customer is being offered when a lab promises not to keep something. It is added as an open problem rather than resolved, because nothing read resolves it and this is a live dispute between two parties, reported at second hand.
State of the Art (as of 2026-09-05)
Anthropic shipped the change, and it went further than the reported plan. On 2026-09-01 Anthropic announced Enterprise Frontier Safeguards (EFS): enterprise customers keep zero data retention while automated misuse detection runs over activity logs stored in cloud infrastructure the customer controls — Amazon S3, Azure Blob Storage or Google Cloud Storage were named — under the customer's own encryption keys and access policies. Automated safety monitoring still scans for misuse; no Anthropic human review is required. Anthropic will not charge for it. Phased rollout from later in fall 2026 (source).
This is the reversal, and the shape of it is not what 2026-08-20 predicted. Bloomberg's reported plan kept the 30-day requirement intact and moved only the storage. What was announced is described across coverage as restoring zero data retention itself for Mythos-class traffic, with two outlets framing it as the retention requirement being dropped for Mythos and Claude Fable 5.1 after enterprise pushback. Anthropic's stated rationale is unchanged — Mythos-class models carry potential for both misuse and autonomous misbehaviour — so the premise this page tracks survived and only the mechanism moved.
What is still not stated, and it is Open Problem 4 verbatim: whether a retention duration still applies to the customer-held logs, or whether the obligation is gone. Nothing read says. The distinction decides whether Anthropic conceded the architecture or only the custody — and this wiki does not have it. One outlet supplies the sharp version of the same gap: Anthropic promises zero data retention, but the customer must verify it worked; verification moves with the data.
Capture note. This is a **+4 day
announcement. www.anthropic.com answers EGRESS_BLOCKED from this pipeline's
sandbox and Anthropic's newsroom has no feed, so it is absent from
state/prefetch.json and reachable only through the daily scrape step. Every
figure above comes from secondary coverage; the first-party post was not read.
This is the third time this wiki has recorded the same failure shape — a Tier-1
source with no feed produces no error when it produces nothing — after the
Alignment Science blog and Meta AI's ai.meta.com.
The convergence is now complete on the mechanism and unresolved on the substance. Ten weeks ago the two labs held opposite positions. Both now propose customer-controlled infrastructure, and neither has published what the lab receives from it. The disagreement this page was created for — must content be retained — has been answered "no" by both parties in practice, without either publishing the evidence that made it answerable.
Anthropic — the change as it was reported on 2026-08-20, at the weakest tier of sourcing this wiki accepts. On 2026-08-20, Bloomberg reported — citing an unnamed source — that Anthropic plans to modify the June policy. Reported terms: enterprise customers still required to retain for 30 days, but given the option to hold that data on their own cloud infrastructure; rollout later this year; developed over months with more than 100 customers, including Salesforce, and with customers in highly regulated industries. Framed across outlets as a response to enterprise backlash (source).
No first-party Anthropic statement was surfaced by targeted search at the time. This wiki's conflict rule ranks official announcement above third-party reporting, and this was third-party reporting of an intention. It was recorded as a reported plan — and it is kept here rather than deleted, because the plan and the announcement differ on the one point that matters: Bloomberg's source said the 30-day requirement stayed. Twelve days later the announcement did not say that. Holding the reported version to the announced one is what makes the difference visible.
Unanswered by anything read, and load-bearing: how a 30-day retention obligation is satisfied on infrastructure the customer controls while still producing the cross-request detection the policy exists for — that is, what Anthropic receives. Also unstated: whether the option restores the voided ZDR agreements or replaces them with a different contractual object, and whether Fable 5's total lack of ZDR support is affected.
Anthropic — retention required (the policy in force 2026-06-09 to 2026-09-01, superseded by EFS above). Since 2026-06-09, prompts and outputs for Mythos-class models (Claude Fable 5, Claude Mythos 5) are retained 30 days for trust and safety, on every surface where the models are offered. It overrides negotiated zero-retention agreements with no opt-out, and Fable 5 does not support ZDR at all. Stated safeguards: no training use, all human access logged, deletion after 30 days "in almost all cases", and human review only through a controlled path after an automated flag. Stated purpose: researching and mitigating jailbreaks — novel attacks, multi-request abuse — and reducing false positives in the safeguard layer (source).
OpenAI — retention claimed unnecessary. On 2026-08-19 OpenAI restated ZDR for eligible API customers (no retention after processing, no personnel review, no training use without opt-in, customer-held encryption keys or customer-controlled infrastructure) and previewed Private Safety Processing: a mechanism that identifies misuse patterns across related interactions and sends OpenAI a narrowly defined safety signal without exposing the underlying prompts or responses. Aleah Houze, Head of Product Policy, gives the rationale — risks at the frontier emerge "when you look over time at multiple interactions", not from a single pair. Enterprise and API only; broader rollout and a technical white paper in September 2026 (source).
What neither has published: any accuracy or false-positive figure for the detection each policy exists to enable. Anthropic states that reducing safeguard false positives is a purpose of retention without giving a rate; OpenAI describes a signal without saying what it contains or how often it fires. This is the third time in two weeks that a safety-relevant classifier has been announced on this wiki with no published error rate — the others being Anthropic's Claude Code auto-mode classifier (89% recall, no false-positive rate) and OpenAI's age-prediction system for ChatGPT for Teens.
Open Problems
- Can a "narrowly defined safety signal" be derived cross-request without retention? The obvious implementations — hashes, embeddings, running summaries, model-side state — are all derived from content and all persist something. Whether the result still satisfies ZDR as customers understand it is the question the September white paper has to answer, and nothing read answers it now.
- Is the disagreement about capability or about capacity to build? Anthropic's policy may reflect a judgement that the private technique does not work, or simply that it had not built one by June. No document read states which.
- What does a customer do about it? Anthropic's requirement is contract-overriding with no opt-out; the option set for a regulated buyer is "accept 30-day retention" or "do not use the model". That is a procurement fact with no published guidance behind it. Partially answered on 2026-08-20 — a third option is reported to be coming, "retain it yourself" — and the answer arrived through Bloomberg rather than through either the policy page or the customer's contract, which is itself the procurement fact. Answered 2026-09-01: the third option shipped as Enterprise Frontier Safeguards, at no charge, and the buyer's option set is now "hold your own logs" rather than "accept retention or leave". The elapsed time is the finding — twelve weeks from a contract-overriding requirement to its withdrawal — and the sequence that produced it was reported backlash, a leaked plan, then a product.
- Does customer-held retention actually change the safety question, or only the custody question? If the lab still needs to compute over the content to detect a cross-request pattern, moving the storage moves liability and residency without moving the capability. If it does not, then content retention was never the requirement and OpenAI's position was right in June. Nothing read says which, and it is now the single most consequential unknown on this page. Still unanswered after 2026-09-01, and now harder to dismiss. EFS keeps automated monitoring while removing Anthropic's custody, which is the second branch — but nothing read states what the monitor extracts, where it runs, or whether a retention duration survives on the customer's side. Anthropic has now built the thing its June policy implied was not buildable, and published no account of what changed. Either the capability existed in June, or it was developed in twelve weeks, or the monitoring is weaker than what retention supported. The three have very different implications and nothing read distinguishes them.
- Does saturation cut the other way? Anthropic reported on 2026-08-14 that its task-based evaluations have saturated and no longer register capability gains (source). If pre-deployment evaluation is losing resolution, post-deployment monitoring carries more of the assurance — which strengthens the case for retention and raises the stakes on whether the private alternative works.
- Nobody has published a comparison. No third party has measured detection quality with content against detection quality with a signal alone, so the central factual dispute between the two labs is currently untested.
- The policy now has a price attached, and it is the first number this page has ever held. Open Problem 3 asked what a customer does about a contract-overriding retention requirement. The Ramp AI Index for August 2026 supplies part of the answer from the demand side: Ramp economist Ara Kharazian names "price + data retention requirements" as the two reasons Claude Fable 5 "disappointed both in adoption and real-world application", against 11.4% of Anthropic dollar spend and 6% of tokens two months after launch (source). Read carefully: this is an analyst's attribution in a spending index, not a customer survey, and it is confounded with price in the same sentence and not separated from it anywhere in what was read — Fable 5 costs twice GPT-5.6 Sol, which alone would predict the shortfall. So the honest statement is that retention has been named as a commercial drag by a party with data, and not yet isolated as one. The distinction matters because the isolated version is the one that would tell a lab what its safety policy costs, and this wiki does not have it. → Model Routing
- Does a retention promise cover de-identified derivatives, and can either side ever check? Raised 2026-09-08 by OpenAI's two-sentence reply to Buckmaster: access denied outright, influence declared unable to be ruled out (source). Every mechanism this page tracks — durations, custody, customer-held keys — governs a log. None of them governs a derivative already folded into weights, and nothing read from any lab states whether a zero-retention commitment reaches that far. The asymmetry is the hard part: an access log can be audited and a training corpus's provenance, at frontier scale, is currently audited by nobody. Until some lab publishes what "de-identified derivative" excludes, a retention guarantee and a training guarantee are different promises that customers have no way to tell apart.
Key Papers
- HarmProfile: Characterizing Harmful Distributions in Frontier LLMs (arXiv:2608.14577) — proposes characterising a model by the distribution of its harmful outputs rather than by a pass/fail rate; the instrument class that a retention-backed monitoring corpus makes possible, and the one that survives a saturated threshold.
- No paper read evaluates privacy-preserving cross-request abuse detection for LLM APIs. This is the gap the September white paper would fill, and the reason it is worth reading when it appears.
Referenced by
Sources
- sources/blogs/buckmaster-alpoge-2026-09-08-fluid-blowup-dispute.md
- sources/blogs/anthropic-2026-09-01-enterprise-frontier-safeguards.md
- sources/blogs/ramp-2026-08-23-ai-index-august-2026.md
- sources/blogs/openai-2026-08-19-zero-data-retention.md
- sources/blogs/anthropic-2026-06-09-mythos-class-data-retention.md
- sources/blogs/anthropic-2026-08-14-risk-report-august-2026.md
- sources/blogs/anthropic-2026-08-20-retention-policy-change.md
- https://www.axios.com/2026/08/19/openai-previews-zero-retention-safety-system-as-anthropic-requires-data-logs