AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/anthropic-2026-05-05-model-spec-midtraining.md

Model Spec Midtraining: Improving How Alignment Training Generalizes

TL;DR

Train the model on documents about its own spec before you train it on examples, and agentic misalignment falls from 54% to 7%. Model Spec Midtraining (MSM) inserts a stage between pre-training and alignment fine-tuning in which the model is trained on synthetic documents discussing its Model Spec. With a spec addressing self-preservation and goal-guarding, Qwen3-32B's agentic misalignment rate goes 54% → 7%, against 14% for a deliberative alignment baseline. Used as an instrument, MSM then shows that explaining the values underlying rules improves generalization, and that specific guidance generalizes better than general guidance (source).

Captured 2026-10-01 at day +148. See Capture below — the lateness is the finding.

Authors & Org

Chloe Li (first author, corresponding), with co-authors not named in anything read. Published by Anthropic as Anthropic Fellows research on the Alignment Science blog, 2026-05-05. Paper: arXiv 2605.02087. Code: github.com/chloeli-15/model_spec_midtraining.

Method

The stated problem is that standard alignment fine-tuning can produce shallow alignment that generalizes poorly, partly because demonstration data underspecifies the desired generalization — a set of examples shows what to do without conveying how far it should reach.

MSM's answer is to put the specification itself into training, as text, before the examples arrive:

StageContent
Pre-trainingordinary corpus
Midtraining (MSM)synthetic documents discussing the model's Model Spec
Alignment fine-tuningdemonstration data
Two things are claimed for it: the model learns the content of the spec, and
the spec shapes how it generalizes from whatever demonstrations follow.

Results

Spec addressing self-preservation and goal-guarding, evaluated on Qwen3-32B:

ConditionAgentic misalignment rate
Baseline54%
Deliberative alignment baseline14%
MSM7%
MSM was then used as a tool to compare specs, rather than only as a training
method. Two findings:
  • Explaining the values underlying rules improves generalization.
  • Specific guidance generalizes better than general guidance.

No Claude result, and no second model. Every figure above is Qwen3-32B's.

Significance

It is a result about generalization, not compliance, and that is the rarer claim. AI Alignment holds several techniques that reduce a measured bad behaviour; this one argues the reason alignment training fails to transfer — underspecified demonstrations — and fixes it upstream of the demonstrations.

It is the training-time counterpart to a proposal this wiki recorded four months later. Safety Cases (2026-09-28) asks for a structured document before a training run continues, read by people. MSM puts a structured document into the run, read by the model. Both treat a written specification as the load-bearing artefact; they disagree about the audience.

The second finding is the one with teeth. "Explaining the values underlying rules" outperforming bare rules is the same shape as Relic: From Multi-Agent Collaboration to Persistent Organizational Capability's result that a rule as a binding beats the same rule as text — except MSM finds that reasons beat rules, where Relic finds mechanism beats documentation. Two 2026 results on the same question pointing in different directions, and this wiki does not resolve it.

Open Questions

  • Whether it is used in production. Nothing read states that Anthropic applies MSM to Claude, and the only numbers are on an open-weight third-party model.
  • What it costs. No compute figure for the midtraining stage.
  • Whether 7% is a floor or a plateau. One spec, one model, one behaviour class.
  • Whether a spec can be over-specified. "Specific beats general" invites the opposite failure and nothing read tests for it.

Capture

This page is 148 days late, and the reason is a check that could not run. agents/daily-run.md requires the Alignment Science blog's article list be checked against sources/ every run. alignment.anthropic.com has answered EGRESS_BLOCKED on every run since it was added — fourteen consecutive as of today — so the check was recorded as attempted and never actually performed.

This run substituted a WebSearch pass for the blocked fetch. It surfaced four posts absent from this wiki, of which this is the oldest and the only one captured today:

PostStatus
Model Spec Midtraining (2026-05-05)this page, +148 days
Poisoning Fine-tuning Datasets of Constitutional Classifiersabsent — title only
Abstractive Red-Teaming of Language Model Characterabsent — title only
Petri 2.0absent — title only (Petri 1.x is held)
The prior measurement of this gap — *"of the four things it published while this
wiki has existed, two were missed entirely"*, SLEIGHT-Bench at +73 and Diffuse AI
Control at +38 — was taken against a list nobody could read. The blog has
published at least ten posts in 2026 and this wiki now holds five.
A source can be listed, cited 24 times, and never fetched; what was new today
is that the list was finally read, by a different tool than the one that keeps
failing.

Cite

Model Spec Midtraining: Improving How Alignment Training Generalizes, Alignment Science Blog, 2026-05-05, https://alignment.anthropic.com/2026/msm/ · arXiv 2605.02087 (source).

Referenced by

Sources