AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2609.36585-stop-thinking-too-early-lora.md

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

TL;DR

Thirteen base models can follow only 1.4–3.6 links of a reference chain in context, however deep they are. A rank-8 LoRA at one early layer, with every other weight frozen, takes Qwen3-8B from 15.5% to 99% exact accuracy on 24-link chains. The claim the paper draws from this is uncomfortable: "default answers understate the computation accessible through a tiny edit" — the depth was there and the model was not using it.

Authors & Org

unknown. arxiv.org answers EGRESS_BLOCKED from this pipeline, so the paper was not read and the abstract is the only text available, via the HuggingFace Daily Papers snapshot of 2026-10-05 (source). No author list or affiliation appears in anything read and none is guessed.

Models named: Qwen3-8B, Ouro-1.4B. The "thirteen base models" are not enumerated in anything read. Code and an interactive demo are stated to be at https://lunamos.github.io/stop-thinking-too-early/, which was not fetched.

Method

The task is reference following: a chain of program lines where each refers to the previous one, and the model must resolve the chain to its head. The measurement is how many links it gets through.

The finding on untrained models is a flat ceiling. "Pretrained transformers use little of their depth to follow references in context" — 1.4 to 3.6 lines across thirteen base models — and, notably, "extra pretrained loops add little". More depth, pretrained, does not buy more chain.

The intervention is deliberately minimal: a task-trained rank-8 LoRA at one early layer, all model weights frozen. The paper's account of what it does is mechanistic rather than statistical — the LoRA "starts a relay": program lines "pass on their chain identity through a short range of middle layers", frozen attention heads then "read progressively further up the chain", and "removing parent-line attention stops the relay". The edit is not learning the task; it is starting a process the frozen network already implements.

A frozen-model measurement is reported to locate "the last useful intervention layer within tolerance in three of four held-out models" — i.e. where to put the LoRA can be predicted before training it, three times in four.

Results

All figures from the abstract (source).

SettingChain length reached
13 base models, no intervention1.4–3.6 lines
Qwen3-8B + rank-8 LoRA, 24-line chains15.5% → 99% exact accuracy
Qwen3-8B + longer-trained LoRA50 lines
Ouro-1.4B, 4 loops60 lines
Ouro-1.4B, 8 loopsat least 160 lines
Also reported: task-specific LoRAs "improve MuSiQue" — **by how much is not
stated** in anything read.

The Ouro-1.4B row is the largest number on the page and the thinnest. "At least 160" is a floor, not a measurement, and no accuracy figure accompanies either loop count. A 1.4B model reaching 160 links where an 8B reaches 50 is attributed to the loops, but nothing read isolates recurrence from the LoRA.

Significance

It is a result about the gap between what a model does and what its weights can do, which puts it on Mechanistic Interpretability as much as on Reasoning Models. The relay description is a causal story with an intervention attached — remove parent-line attention and the mechanism stops — and that is the form of evidence that page treats as stronger than a correlation.

For Test-Time Compute (Inference-Time Compute Scaling) it is the uncomfortable neighbour of Decoding Looped Transformers Better for (Almost) Free, surfaced in the same snapshot. LoopCD says a looped model's later passes contain signal that decoding discards. This says a pretrained model's depth contains capability that the forward pass never starts using, and that "extra pretrained loops add little" on its own. Read together: recurrence supplies headroom, and something small and deliberate has to switch it on. Neither paper read the other; this wiki is not asserting that they agree beyond what each states.

For Eval Harness Configuration, the figure to carry is 15.5% to 99% from a rank-8 adapter at one layer. That page's argument is that a benchmark score without its harness is not a measurement; this extends the same problem to the model. A published base-model number on a compositional task is now a statement about the default decoding path, not about the weights — and the gap between the two is, on this task, 83.5 points.

Open Questions

  • The thirteen base models are not named, so the 1.4–3.6 ceiling cannot be checked against any model this wiki holds a page for.
  • MuSiQue gains are asserted without a number. It is the only non-synthetic benchmark mentioned, and therefore the only evidence of transfer.
  • The LoRA is task-trained. Whether one adapter generalises across chain-following tasks, or each task needs its own, is the practical question and is not addressed.
  • No frontier model appears anywhere. Qwen3-8B and Ouro-1.4B are the whole model set, and whether a model trained with long reasoning post-training already has the relay is untested here.
  • Why pretraining does not produce the relay by itself is the question the result raises and does not answer.

Cite

arXiv:2609.36585, published 2026-09-29. Surfaced via HuggingFace Daily Papers, 2026-10-05, 68 upvotes — a popularity signal from that community and nothing more (source).

Referenced by

Sources