AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/test-time-compute.md

Test-Time Compute (Inference-Time Compute Scaling)

Definition

Any technique that improves output quality by allocating additional compute at inference time. The trained model stays fixed, but it thinks/searches/verifies more during the process of producing an answer.

Representative forms:

  • Chain-of-Thought (CoT) — explicitly generating intermediate steps
  • Self-consistency — sampling multiple answers and taking a majority vote
  • Tree-of-Thoughts / search — branching exploration
  • Verifier-guided search — a separate verifier evaluates candidates
  • Reflection / critic — self-criticism followed by retry
  • Tool augmentation — calling external calculators/search

Why It Matters

  • Matches personal interests (reasoning 1.3x)
  • One of the biggest paradigm shifts of 2024-2026 — beyond "bigger models," toward "models that think longer"
  • The core inference-cost trade-off, assuming training-cost limits have been reached
  • The operating mechanism behind Reasoning Models

State of the Art (2026-07)

Two months on, the interesting question is no longer how much inference compute helps but where it is spent and whether it is the compute you think:

  • Horizon, not parameters — Agents-A1 (2026-06) matches trillion-parameter models on agentic benchmarks with 35B active by scaling the agent horizon — trajectory length and ability breadth — rather than model size. Test-time compute spent on a longer trajectory bought what 30× the parameters did
  • The compute may not reach the policy you trained — the MIPI paper (2026-06) finds RL optimising a training proxy while a different engine serves inference, so effort spent at inference is applied to a policy the training never targeted
  • It became a price tier — Gemini 3.5 Pro puts "Deep Think" behind the $250/month Ultra tier. Where test-time compute was a research knob it is now a line item, which is its own answer to "how far is it efficient"
  • Failure scales too — Ring-Zero documents behaviours appearing only above ~100B, including stalling on ambiguous context. More inference compute does not monotonically buy more reliability → Reasoning Models

2026-10-04 — a seventh mechanism, and it sits in the gap between the model and the harness

Every mechanism on this page spends compute inside a generation (longer chains, more samples of a whole answer, more forward passes) or across trajectories. Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents spends it on the single next action: sample several candidate shell commands, verify them, forward one for execution, and leave the generator and the harness untouched (source).

On TerminalBench-Lite with a TMAX-9B generator, Pass@1 goes from 50.00% for the base agent to 68.03% with 8 sampled actions verified by a GPT-5.6 Sol verifier.

The conditional is the result, and it is the most useful sentence on this page for anyone allocating a budget. More action sampling "yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator." The better actions were already in the generator's distribution; the compute that paid off was spent on telling them apart, not on producing more of them. That makes this a direct instance of the ## Open Problems entry below on verifier accuracy becoming the new ceiling — here the ceiling is measured rather than asserted.

Two further findings bear on allocation:

  • Action scaling and trajectory scaling combine to beat trajectory scaling alone, at lower estimated token cost. This is the first result on this page comparing two allocation axes against each other on cost rather than on score.
  • Distilling the strong verifier into TMAX-9B improves Pass@1 further, with the action generator unchanged — so the dependency on a frontier verifier is at least partly convertible into weights.

What this does not establish. The 68.03% configuration calls a frontier model on every action and no cost accounting for that is given — the token-cost comparison reported is action-scaling versus trajectory-scaling, not either against the base agent. TerminalBench-Lite is the only benchmark named; "additional models, benchmarks, and harnesses" is asserted without specifics. And TMAX-9B is not a model this wiki holds a page for, so the 50.00% baseline is uncalibrated. The paper was not read — arxiv.org is blocked from this pipeline and the abstract is the only text available.

The awkward corollary belongs here rather than being left to Eval Harness Configuration: an 18-point Pass@1 swing with the model and the harness both unchanged means a TerminalBench-Lite figure is underdetermined unless the action-selection policy is stated.

2026-08-25 — the sixth mechanism, and it is about taking compute back

Every mechanism below answers "where should more inference compute go". ParaTempo: Efficient Parallel Reasoning via Temporal Confidence (arXiv:2608.16425) asks the reverse question of parallel reasoning — which branches have already decided, and can therefore stop — and governs the entire budget from one signal (source).

The signal is the contribution. It names the three existing control signals and why each fails: final-answer consensus is delayed (it arrives after the compute is spent), local token confidence is weakly tied to reasoning progress, and isolated intermediate probes are too noisy for branch-level control. Temporal confidence is the third one denoised by aggregation — each branch is probed periodically for a tentative answer distribution, and the measure is how sharply the recent probes concentrate, a trend rather than a reading.

One number, four actions, and no synchronization between trajectories: prune low-confidence branches, retire branches that persistently commit, fork new branches with the freed compute, and stop globally when the confidence-weighted vote concentrates. Retiring a confident branch is the counterintuitive move — it treats commitment as completion rather than as a winner worth extending, and it is what turns a pruning heuristic into a reallocator.

Reported: latency down 21.8–32.2%, tokens down 18.1–30.3%, accuracy "competitive" — with no accuracy figure, no named benchmark, and no statement of whether the probing cost is netted out. Training-free, like the day's two scaffold-evolution results on Agents (LLM Agents).

The unexamined risk is structural: the global stop fires when the confidence-weighted vote concentrates, which accelerates exactly the convergence that self-consistency relies on being slow enough to be informative. Nothing read tests the case where the majority converges on a wrong answer.

2026-08-24 — the mechanism crosses into robot control, and lands on the sequencing decision

Every mechanism catalogued below spends inference compute on text reasoning. τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation (arXiv:2608.16885) moves the idea into a hierarchical vision-language-action model, and the interesting part is where it puts it: not in the motor policy, but in the high-level decision of which subtask to attempt next. The paper's observation is that standard hierarchical VLAs make each such choice in a single forward pass, which leaves "no mechanism to allocate additional computation to difficult or consequential choices" — the same complaint that produced adaptive-depth and search-based decoding in the text setting, arriving independently in robotics.

The construction is search-with-a-verifier with the verifier swapped: the policy uses execution memory to propose a subtask and, when needed, searches over alternatives against a world model before committing. Trained on 40,115 hours of heterogeneous real-world data. Reported outcome: additional test-time compute substantially improves next-subtask prediction accuracy in-domain and under distribution shift, and those gains translate into higher closed-loop success on long-horizon manipulation (source).

Two things this page should hold against it. The abstract publishes no numbers at all — no accuracy, no success rate, no compute-versus-performance curve — so "compute-scalable" is so far a design property and not a measured one, and this page has no figure to put beside the text-domain mechanisms below. And the verifier is unaudited: Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning (arXiv:2608.18746) established two days later that a latent world model can probe well and still rank plans wrongly, which is exactly the failure a search step would silently inherit. τ_0-VLA reports no diagnostic pointed at its own world model.

It also runs against Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence (arXiv:2608.16590) from the same week, whose premise is that physical control needs decisions faster than a large model can make them. The likely reconciliation — Zetta is talking about the control loop, τ_0-VLA about the subtask boundary above it — is not stated by either, and where the boundary falls is unmeasured.

2026-08-23 — the prefill side of the bill, taken to production

Every mechanism on this page spends compute at inference; none of them has said what the baseline costs. FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving (arXiv:2608.19758) works the other end — the prefill phase, where attention's quadratic term bites hardest — and takes a block-sparse prefill attention from prototype to something a serving stack runs: a mean correction term holding approximation error down at extreme sparsity, the operator rewritten to match FlashAttention-3/4 with FP8, and native paged KV cache and continuous batching so it can back a framework such as SGLang.

On NVIDIA H20 at 128K context: up to 47.26× over FlashAttention-2 in FP8 and 27.19× in BF16, and — the figure that actually bears weight — 30.49× in FP8 against an FA3/4-aligned dense baseline (source).

The relevance here is the budget, not the method. Adaptive-depth, beam-search and knowing-when-to-quit all assume a context the serving stack can afford to fill; this is what makes that assumption cheap. And it lands in the same fortnight that MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use (arXiv:2608.20202) and SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? (arXiv:2608.19799) argue the context should often be smaller — so the constraint this removes may not be the binding one.

No accuracy cost reached this capture. A sparse-attention speedup without its quality trade stated is half a result, and "up to" carries every figure above.

2026-08-21 — a fifth mechanism, and it spends the compute inside the forward pass

Every mechanism above spends inference compute on tokens — longer chains, search, verifiers, or (in the full-bandwidth transformer's case) a widened channel between decoding steps. Looped Language Models Improve Compositional Tool Calling (arXiv:2608.18171) spends it on recurrent depth: a looped model reuses the same parameters over several passes, so depth varies at inference time without adding parameters, and the paper tests it in the middle of an agent's tool-use workflow rather than on a reasoning benchmark.

The boundary it draws is the useful part. Recurrent computation helps compositional and dependency-aware tool use — coordinating multiple API calls, holding intermediate state, preserving dependencies — while gains on isolated API invocation are "smaller and more model-dependent". A single call is close to a lookup; composing calls requires carrying state across steps, and only the second buys anything from extra depth. Evaluated on API-Bank, BFCL and NESTful, with looped and non-looped models trained under matched SFT recipes.

And the allocation argument arrives for the third distinct mechanism this month. Accuracy generally rises with recurrent depth, but adaptive inference gets a better compute–performance trade-off by spending only when needed — the same finding as Thought-Level Beam Search for Reasoning (arXiv:2608.08020)'s reframing of test-time compute as allocation rather than budget, and as Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211)'s argument that the right budget is sometimes zero. Three unrelated mechanisms, one conclusion: where to spend beats how much.

The limitation is severe and worth stating plainly: the abstract carries no numbers at all. "Generally benefits" and "smaller and more model-dependent" are directions, not magnitudes, so nothing here says whether looping beats simply using a larger model at matched FLOPs — the obvious alternative use of the same compute, and the comparison the paper does not report.

2026-08-19 — a fourth mechanism, and it needs no new method at all

The three mechanisms below change what a budget buys. R³-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets (arXiv:2608.16033) changes nothing and shows that models already underuse the budget they have (source).

Its design is what gives it force. Six problems share one budget across mathematics, competitive programming and abstract reasoning, in tool-free and agentic settings. The comparison is not against a stronger model — it is against an offline oracle built from the same model's measured single-problem response curves. So the model is being compared to itself under a better allocation policy.

FindingAs reported
Oracle vs contest, six modelsoracle mean matches or exceeds contest mean in all 72 cells; strictly higher in 71
Equal-allocation replay, moderate tool-free pressureexceeds contest performance for four of six models
Fixed schedulers, strong agentic pressureat least one exceeds the contest mean in six of nine cells
Strategy updatinglimited; failure patterns are pressure-dependent; no policy dominates across domains
The sentence to keep is the second row. Splitting the budget six ways with no
reasoning at all beats the model's own reasoning about how to split it, for four of
six models. The other three mechanisms all propose machinery; this baseline is free.

It converges with the same-day negative result on Eval Harness Configuration. How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905) names the faculty missing across all 8 harness-model combinations it tested as a metacognitive loop — checking output against evidence, revising, questioning the path. Allocating a shared budget requires judging your own likely success on each problem before spending, which is that faculty priced in tokens. Neither paper cites the other; the pairing is this wiki's and is labelled as such. It also restates Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211)'s systematic miscalibration from the allocation side rather than the refusal side.

Consequence for how this wiki reads a reasoning score. Every such figure here comes from independent per-task budgets. R³-Bench says that regime overstates what the same model does when it must allocate — the condition an agent actually runs under. Nothing read quantifies the offset, and no source attempts one, so no figure on any model page is adjusted; this is recorded as a known bias in the measurement, not as a correction to apply.

Unresolved and material: the abstract gives a direction and never a magnitude ("strictly higher in 71 of 72" is satisfied by a uniform 0.3 points or by 15), and the budget's unit is not stated — which makes it incomparable to the token-count claims in the mechanisms below, the same defect Eval Harness Configuration tracks.

2026-08-18 — a third mechanism, and it argues the right budget is sometimes zero

Yesterday's two mechanisms both assume the model will eventually answer: Thought-Level Beam Search for Reasoning (arXiv:2608.08020) moves compute between live traces, Full-bandwidth transformer (arXiv:2608.08888) reduces what a step needs. Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211) is categorically different — it says a substantial share of inference spend goes on questions where no amount of reasoning produces an answer, and the model does not know it (source).

Its diagnosis, stated without figures: universal capability overreach, systematic miscalibration between capability and behaviour, and a dominant failure mode of specious reasoning — output that looks valid and contains subtle errors — which escalates with task difficulty. Its fix, CaRL, shapes reward to incentivise refusal over futile reasoning and converts past failures into refusal supervision.

MechanismWhat it changesAnswer still produced?
Thought-Level Beam Search for Reasoning (arXiv:2608.08020)allocation across live tracesyes
Full-bandwidth transformer (arXiv:2608.08888)what one decoding step needsyes
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211)whether to start at allno — refusal is the output
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning (arXiv:2608.14277)where a student learns when to stopyes
The fourth row is the day's quieter finding. SimpleOPD masks the advantages of
termination tokens — </think>, `<im_end>` — because a short-context student
distilling from a long-context teacher inherits the teacher's stopping behaviour
and blows up its own response length. **Stopping is a learned behaviour that
transfers badly**, which is the same phenomenon CaRL treats as miscalibration,
observed inside a distillation pipeline instead of at inference
(source).

The ceiling is not additive. Four mechanisms all claiming to remove wasted tokens compose only if the tokens they remove are disjoint, and nothing read establishes that any two of them are. None of the four cites another.

The caution this page records against CaRL specifically: an objective that pays a model to refuse has an obvious degenerate solution, and the abstract's only defence is "without sacrificing utility" with no figure attached. A refusal rate is the single number needed to read the result, and it is not published.

2026-08-17 — the waste got measured, from both ends, on the same day

This page's open problem has been "how much inference cost can users absorb". Two things arriving together reframe it: a large share of that cost may be spent on trajectories that were already lost, in which case the question is not absorption but waste.

The cost, measured by a practitioner. Running Qwen 3.8 27B locally as a 17GB GGUF on an M5 Max MacBook Pro, Simon Willison recorded a single "generate an SVG" prompt consuming 22,276 reasoning tokens against 3,223 output tokens — 6.9 reasoning tokens per output token — in 21 minutes. That is the behaviour of the model's shipped default: Qwen3.8 ships reasoning_effort at xhigh, with medium and low available and not chosen for the user. His own summary is that the model "will happily burn 20,000 reasoning tokens on a two-sentence answer" (source).

The recovery, claimed by an algorithm. Thought-Level Beam Search for Reasoning (arXiv:2608.08020) argues the critical question has shifted "from how much compute to spend, to where to allocate it", and reports up to 68.5% fewer total tokens than standard parallel sampling with accuracy up — +6.7 absolute on HMMT-24, +3.3 on AIME-25 — and >2× throughput, under identical hardware (source).

The two are not about the same model and neither cites the other; the pairing is this wiki's. But a token count and a claimed recovery rate have never sat on this page together before, and they point the same way: the marginal reasoning token is not the one being priced.

A third result in the same batch attacks the same waste before inference. Full-bandwidth transformer (arXiv:2608.08888) feeds the previous top-layer hidden state back into the next decoding step, and reports shorter reasoning traces at equal or better accuracy — the implication being that part of a chain of thought exists to carry computation between steps, not to perform it, and that a wider vertical channel makes those tokens unnecessary (source).

The consequence for reading a Pricing cell. Three mechanisms — reallocation between traces, a wider latent channel, and a default effort level — each change the token count for the same answer without changing the price per token. A model page's Pricing row is a rate; none of these is visible in it. That is the same observation Eval Harness Configuration makes about benchmark scores, arriving at the invoice instead of the leaderboard.

Scaling Laws (open question)

  • Does test-time compute follow power-law scaling like training compute?
  • Where is the saturation point?
  • Does scaling differ by domain (math vs coding vs science vs real-world)?

→ Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling provides the most recent data in this area

A training-side scaling law arrived on 2026-08-25 and it is about the search, not the model. Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts (arXiv:2608.20061) predicts the optimal learning rate for large MoE training in two transfers — a μP adaptation under which the optimum holds across model width, then a linear regression on small-proxy optima that extrapolates along the token axis to 10 trillion tokens at R² = 0.95, validated by pretraining a 155B total / 17B active model from scratch. It answers none of the three questions above, and it is filed here because it is the only quantity in this area with a published fit statistic: the scaling law that currently works is one about what a lab must spend to configure a run, not about where capability comes from (source)

A second training-side law arrived on 2026-09-30, and this one does answer a question above — for recurrence. Scaling Laws for Looped Mixture of Experts (arXiv 2609.40316, HuggingFace Daily Papers 2026-10-02, 15 upvotes) is described as the first scaling law to model recurrence and sparsity jointly alongside model size and data, recovering the standard dense and MoE laws as special cases. Reported efficiencies: sparsity delivers ~3× active-parameter efficiency, recurrence ~2× total-parameter efficiency on reasoning, and at matched training compute a looped MoE with law-derived recurrence matches a ~2× larger non-looped MoE on reasoning benchmarks — "while enabling test-time scaling through recurrence" (source).

That last clause is why it is filed here and not only under training: recurrence is a depth knob that can be turned at inference, which makes it a test-time mechanism chosen at architecture time. The paper was not read — arxiv.org is blocked from this pipeline — so this is the abstract only, and no fit statistic is quoted, unlike the 2026-08-25 law above. It does not address saturation or domain dependence.

The ~3× sparsity figure is a general claim and is not evidence for Kolibri-1's "four times its active parameter count", which is a different model, a different measurement and a vendor's own table. They are the same question asked twice, and this page does not treat either as support for the other.

2026-10-05 — the first mechanism here that spends nothing, and then hands compute back

Every other mechanism on this page buys something: a longer chain, more samples, more trajectories, a verifier on each action. LoopCD buys nothing and the compute curve moves down.

Decoding Looped Transformers Better for (Almost) Free (arXiv 2610.02185, HuggingFace Daily Papers 2026-10-05, 38 upvotes) observes that a looped Transformer already produces a decodable prediction at every recurrent pass, and that standard decoding discards all but the last. Because "earlier loops embody less computation", recurrence "inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training" — so contrastive decoding needs no second model. Training-free, in two variants: LoopCD-Logits (one extra output pass) and LoopCD-Hidden (zero output overhead).

ModelBenchmarkBaselineLoopCD
Ouro-2.6B-ThinkingAIME 2024 pass@161.88%73.33%
HuginnHumanEval pass@122.56%31.71%
And then the refund: the gains *"enable halving the number of recurrent loops while
still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by
22.5% to 48.2%"*
(source).

Read it against the 2026-09-30 scaling law in the Scaling Laws (open question) section above. Scaling Laws for Looped Mixture of Experts argues recurrence is worth buying — ~2× total-parameter efficiency on reasoning — and that is why recurrence is filed on this page at all. This paper argues that part of what was bought is already paid for and was being thrown away at the output. Both are abstract-only reads of papers that do not cite each other, and this wiki is not asserting a joint result; the two are recorded as the same knob approached from opposite ends, three days apart.

What is not established: the "four looped Transformer families" are unnamed; each model carries exactly one benchmark, so neither variant is shown to deliver both gains; which earlier pass is contrasted against is unspecified; and the 22.5–48.2% FLOPs range is unattributed to any model or benchmark. No looped model in this wiki has a page, and nothing read tests the method above ~2.6B parameters.

The same snapshot's other depth result, and it cuts the other way

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It (arXiv 2609.36585, same 2026-10-05 snapshot, 68 upvotes) reports that thirteen pretrained base models follow only 1.4–3.6 links of an in-context reference chain, and — the clause that belongs on this page — that "extra pretrained loops add little". A rank-8 LoRA at one early layer with all weights frozen then takes Qwen3-8B from 15.5% to 99% exact accuracy on 24-link chains, and Ouro-1.4B to 60 links after four loops and at least 160 after eight (source).

So recurrence supplies headroom the default forward pass does not use, and something small has to start it — the paper's own phrasing is that the LoRA "starts a relay". That is a limit on the mechanism this page tracks, not a result about allocation: spending more depth is not the same as using it. The Ouro-1.4B figures carry no accuracy number and "at least 160" is a floor, so they are recorded as reported rather than as a measurement.

Open Problems

  • Cost-quality trade-off — how much inference cost can users/products absorb?
  • Latency — models that think more at inference respond slower. Which use cases can tolerate it?
  • Verifier accuracy — the verifier's own limits become the new ceiling
  • Compute allocation optimization — how to distribute the same budget across step / sample / search?

Key Papers

  • Reasoning Models — the class of models that use this mechanism
  • scaling-laws — separate page TBD
  • Agents (LLM Agents) — multi-step agents are also a form of test-time compute

Open Debates

Notable Statements

  • (Sam Altman): the "automated AI research intern by 2026-09" goal presupposes strong test-time compute + an agentic loop

Referenced by

2026-W40Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified ScalingAgentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310)Agents (LLM Agents)AI AlignmentAI for MathematicsAleph AlphaAn Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsBDH-CQ: In-Context Learning with Recurrent Latent Reasoning (arXiv:2608.09888)Beyond Solver Verdicts: Generative Reward Models for AutoformalizationDecoding Looped Transformers Better for (Almost) FreeDoes On-Policy Distillation Really Distill? From Noisy Teacher to Self-ImprovementEval Harness ConfigurationFlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving (arXiv:2608.19758)Full-bandwidth transformer (arXiv:2608.08888)Gemini 3.5 ProGemini 3.8 Live Extended ThinkingHow Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905)InklingIntern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (arXiv:2608.14290)Jason WeiKnowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211)Kolibri-1Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts (arXiv:2608.20061)Looped Language Models Improve Compositional Tool Calling (arXiv:2608.18171)Mechanistic InterpretabilityMeta^n: Recursive Self-Improvement through Emergent Depth (arXiv:2608.24735)Mid-Harness: Scaling Actions Between Model and Harness for Terminal AgentsModel RoutingOn-Policy or Off-Policy Learning? A Systematic Study of Distillation DynamicsParaTempo: Efficient Parallel Reasoning via Temporal Confidence (arXiv:2608.16425)PARSER: Read in Parallel, Reason in Depth for Long-Context LLM AgentsPost-Training ScalingPrime Agent: A Self-Improving RLM Harness (arXiv:2608.23552)PrismMLQwen 3.8 27BR³-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets (arXiv:2608.16033)Reasoning ModelsRing-Zero: Scaling Zero RL to a Trillion Parameters for Emergent ReasoningRound-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors (arXiv:2608.00675)Sam AltmanScaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B AgentSimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning (arXiv:2608.14277)SMELT: Scaling Laws for Compute-Matched MoE Looped TransformersStateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (arXiv:2608.15089)The Embedder's Dilemma: LLMs Are Better, but at What Cost? (arXiv:2608.12875)The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement LearningThe Tasteful Agent: Measuring and Improving Taste in Long-Horizon TasksThought-Level Beam Search for Reasoning (arXiv:2608.08020)Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes ItTTPO: Test-Time Policy Optimization (arXiv:2608.27448)Unlocking Lossless Speedups in LLMs via Discrete DiffusionWhen Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token AnalysisYour Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMsτ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation (arXiv:2608.16885)

Sources