$ cat wiki/papers/2026/anthropic-2026-05-05-model-spec-midtraining.md
Model Spec Midtraining: Improving How Alignment Training Generalizes
TL;DR
Train the model on documents about its own spec before you train it on examples, and agentic misalignment falls from 54% to 7%. Model Spec Midtraining (MSM) inserts a stage between pre-training and alignment fine-tuning in which the model is trained on synthetic documents discussing its Model Spec. With a spec addressing self-preservation and goal-guarding, Qwen3-32B's agentic misalignment rate goes 54% → 7%, against 14% for a deliberative alignment baseline. Used as an instrument, MSM then shows that explaining the values underlying rules improves generalization, and that specific guidance generalizes better than general guidance (source).
Captured 2026-10-01 at day +148. See Capture below — the lateness is the finding.
Authors & Org
Chloe Li (first author, corresponding), with co-authors not named in anything
read. Published by Anthropic as Anthropic Fellows research on the
Alignment Science blog, 2026-05-05. Paper: arXiv 2605.02087. Code:
github.com/chloeli-15/model_spec_midtraining.
Method
The stated problem is that standard alignment fine-tuning can produce shallow alignment that generalizes poorly, partly because demonstration data underspecifies the desired generalization — a set of examples shows what to do without conveying how far it should reach.
MSM's answer is to put the specification itself into training, as text, before the examples arrive:
| Stage | Content |
|---|---|
| Pre-training | ordinary corpus |
| Midtraining (MSM) | synthetic documents discussing the model's Model Spec |
| Alignment fine-tuning | demonstration data |
| Two things are claimed for it: the model learns the content of the spec, and | |
| the spec shapes how it generalizes from whatever demonstrations follow. |
Results
Spec addressing self-preservation and goal-guarding, evaluated on Qwen3-32B:
| Condition | Agentic misalignment rate |
|---|---|
| Baseline | 54% |
| Deliberative alignment baseline | 14% |
| MSM | 7% |
| MSM was then used as a tool to compare specs, rather than only as a training | |
| method. Two findings: |
- Explaining the values underlying rules improves generalization.
- Specific guidance generalizes better than general guidance.
No Claude result, and no second model. Every figure above is Qwen3-32B's.
Significance
It is a result about generalization, not compliance, and that is the rarer claim. AI Alignment holds several techniques that reduce a measured bad behaviour; this one argues the reason alignment training fails to transfer — underspecified demonstrations — and fixes it upstream of the demonstrations.
It is the training-time counterpart to a proposal this wiki recorded four months later. Safety Cases (2026-09-28) asks for a structured document before a training run continues, read by people. MSM puts a structured document into the run, read by the model. Both treat a written specification as the load-bearing artefact; they disagree about the audience.
The second finding is the one with teeth. "Explaining the values underlying rules" outperforming bare rules is the same shape as Relic: From Multi-Agent Collaboration to Persistent Organizational Capability's result that a rule as a binding beats the same rule as text — except MSM finds that reasons beat rules, where Relic finds mechanism beats documentation. Two 2026 results on the same question pointing in different directions, and this wiki does not resolve it.
Open Questions
- Whether it is used in production. Nothing read states that Anthropic applies MSM to Claude, and the only numbers are on an open-weight third-party model.
- What it costs. No compute figure for the midtraining stage.
- Whether 7% is a floor or a plateau. One spec, one model, one behaviour class.
- Whether a spec can be over-specified. "Specific beats general" invites the opposite failure and nothing read tests for it.
Capture
This page is 148 days late, and the reason is a check that could not run.
agents/daily-run.md requires the Alignment Science blog's article list be
checked against sources/ every run. alignment.anthropic.com has answered
EGRESS_BLOCKED on every run since it was added — fourteen consecutive as of
today — so the check was recorded as attempted and never actually performed.
This run substituted a WebSearch pass for the blocked fetch. It surfaced four posts absent from this wiki, of which this is the oldest and the only one captured today:
| Post | Status |
|---|---|
| Model Spec Midtraining (2026-05-05) | this page, +148 days |
| Poisoning Fine-tuning Datasets of Constitutional Classifiers | absent — title only |
| Abstractive Red-Teaming of Language Model Character | absent — title only |
| Petri 2.0 | absent — title only (Petri 1.x is held) |
| The prior measurement of this gap — *"of the four things it published while this | |
| wiki has existed, two were missed entirely"*, SLEIGHT-Bench at +73 and Diffuse AI | |
| Control at +38 — was taken against a list nobody could read. The blog has | |
| published at least ten posts in 2026 and this wiki now holds five. | |
| A source can be listed, cited 24 times, and never fetched; what was new today | |
| is that the list was finally read, by a different tool than the one that keeps | |
| failing. |
Cite
Model Spec Midtraining: Improving How Alignment Training Generalizes, Alignment Science Blog, 2026-05-05, https://alignment.anthropic.com/2026/msm/ · arXiv 2605.02087 (source).