$ cat wiki/papers/2026/2610.02185-loopcd-looped-decoding.md
Decoding Looped Transformers Better for (Almost) Free
TL;DR
A looped Transformer already computes a decodable prediction at every recurrent pass, and standard decoding throws all but the last away. LoopCD contrasts the final pass against an earlier one to steer token choice — no training, no auxiliary model — and reports AIME 2024 pass@1 61.88% → 73.33% on Ouro-2.6B-Thinking. The second finding is the one that matters more: the gain lets you halve the loops and still match full-depth baselines, cutting forward FLOPs 22.5–48.2%.
Authors & Org
unknown. arxiv.org answers EGRESS_BLOCKED from this pipeline, so the paper was
not read and the abstract is the only text available, via the HuggingFace Daily
Papers snapshot of 2026-10-05
(source).
No author list or affiliation appears in anything read and none is guessed.
Models named: Ouro-2.6B-Thinking and Huginn, plus "four looped Transformer families" which are not enumerated in anything read.
Method
A looped Transformer runs one shared block repeatedly, so parameter count stays low while effective depth rises. The paper's observation is that each loop leaves behind an intermediate representation that is "decodable for the same next token" — the model has, for free, a sequence of guesses at the same prediction, ordered by how much computation went into them.
That ordering is the whole mechanism. Because "earlier loops embody less computation", recurrence "inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training". Contrastive decoding normally needs a deliberately weaker second model to contrast against; here the weaker model is the same model, earlier.
LoopCD is training-free and comes in two variants:
| Variant | Where the contrast happens | Overhead |
|---|---|---|
| LoopCD-Logits | logit space | one extra output pass |
| LoopCD-Hidden | hidden-state space | zero output overhead |
Results
All figures from the abstract (source).
| Model | Benchmark | Baseline | LoopCD | Variant |
|---|---|---|---|---|
| Ouro-2.6B-Thinking | AIME 2024 pass@1 | 61.88% | 73.33% | LoopCD-Logits |
| Huginn | HumanEval pass@1 | 22.56% | 31.71% | LoopCD-Hidden |
| Reported as "substantial, consistent gains at full recurrent depth" across **four | ||||
| looped Transformer families**. |
The compute result is stated separately: the gains "enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%".
Note what the two tables are not. The 11.45-point AIME gain and the 9.15-point HumanEval gain are each reported for a different model with a different variant, so the abstract does not establish that either variant delivers both. No per-variant comparison on a single model appears in anything read.
Significance
It is the first entry on Test-Time Compute (Inference-Time Compute Scaling) where recurrent depth is spent and then partly refunded. That page recorded on 2026-10-04 that Scaling Laws for Looped Mixture of Experts (arXiv 2609.40316) makes recurrence a depth knob turnable at inference, with ~2× total-parameter efficiency on reasoning. This paper turns the same knob the other way: it extracts more from the loops already run, and uses that surplus to run fewer of them. Two papers three days apart, one arguing recurrence is worth buying and one arguing part of it is already paid for.
It also sits oddly against the mechanism this page has been accumulating. Nearly every other entry spends more at inference — longer chains, more samples, more trajectories, a verifier on each action (Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents). LoopCD spends nothing and the compute curve moves down, which makes it a different kind of claim: not a better allocation of a test-time budget, but a reading of signal that the forward pass was discarding.
For Eval Harness Configuration it is another entry in the pattern that page exists for. A looped model's benchmark number is now underdetermined unless the decoding rule and the loop count are stated: the same weights read two ways differ by 11.45 points on AIME 2024, and halving the loops is reported as roughly free. A published Ouro or Huginn score with neither figure attached cannot be compared to one of these.
Open Questions
- The four looped families are not named. "Consistent gains" across them cannot be checked, and Ouro and Huginn are the only two models in any reported number.
- One benchmark per model. AIME 2024 for Ouro, HumanEval for Huginn — nothing read reports both models on both, so whether the gain is reasoning-specific or general is open.
- Which earlier pass to contrast against is not stated. "An earlier recurrent pass" leaves the choice unspecified, and the method's sensitivity to it is the obvious ablation; none is reported.
- The FLOPs range is wide — 22.5% to 48.2% — and unattributed. Nothing read says which model or benchmark sits at either end.
- No looped frontier model exists in this wiki. Ouro-2.6B and Huginn are small research models; whether this transfers to the scale at which Reasoning Models figures are published is untested here.
Cite
arXiv:2610.02185, published 2026-10-01. Surfaced via HuggingFace Daily Papers, 2026-10-05, 38 upvotes — a popularity signal from that community and nothing more (source).