$ cat wiki/papers/2026/2609.39982-mid-harness.md
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
TL;DR
Spend test-time compute at the boundary between the model and the harness: sample several candidate shell actions, verify them, and forward only one for execution. On TerminalBench-Lite this lifts Pass@1 from 50.00% to 68.03% — without changing the generator or the harness.
Authors & Org
unknown. arxiv.org answers EGRESS_BLOCKED from this pipeline, so the paper was
not read and the abstract is the only text available, via the HuggingFace Daily
Papers snapshot of 2026-10-02
(source). No author
list or affiliation appears in anything read and none is guessed.
Models named: TMAX-9B as generator, GPT-5.6 Sol as the strong verifier.
Method
The premise is a distinction this wiki has not previously had a page for: being able to generate a good action is not the same as reliably executing one. The abstract's example is concrete — a wrong package install changes the environment in ways that block later progress, "even when the model could generate a better alternative".
Mid-Harness therefore inserts a step between generation and execution: it samples candidate actions and verifies them before forwarding one. Crucially, the generator and the harness are left unchanged — this is an addition at the seam, not a replacement of either side.
Three things are varied:
- Number of sampled actions (up to 8 reported)
- Who verifies — a stronger external model (GPT-5.6 Sol) or the generator itself (TMAX-9B)
- Verification mechanism — among those evaluated, pairwise verification performs best when TMAX-9B verifies its own candidates
A distillation variant is also reported: responses from the stronger verifier are distilled into TMAX-9B, leaving the action generator unchanged.
Results
All figures from the abstract (source).
| Condition (TerminalBench-Lite, TMAX-9B generator) | Pass@1 |
|---|---|
| Base agent | 50.00% |
| + 8 sampled actions, GPT-5.6 Sol verifier | 68.03% |
| The conditional is the finding, not the 18-point gain. Stated directly: more | |
| action sampling "yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator." The candidates | |
| were already there; what was missing was something able to tell them apart. Sampling | |
| without a good verifier buys nothing. |
Further results:
- With TMAX-9B as its own verifier, pairwise verification is the best of the mechanisms evaluated.
- Distilling the stronger verifier's responses into TMAX-9B improves Pass@1 further, with the generator untouched.
- Combining action scaling with trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone.
- The method "also improves performance across additional models, benchmarks, and harnesses" — none of which are named in anything read.
Significance
It identifies a place to spend test-time compute that Test-Time Compute (Inference-Time Compute Scaling) did not have. That page tracks compute spent inside a generation (longer chains, more samples of a whole answer) and across trajectories. This is neither: it is compute spent on the single next action, at the harness boundary, and the paper's claim is that this is the cheaper axis — higher success at lower estimated token cost than more trajectories.
For Eval Harness Configuration it cuts the other way, and more awkwardly. The harness is unchanged, the model is unchanged, and Pass@1 moves 18 points. That is an 18-point swing attributable entirely to scaffolding between the two, which is precisely the quantity that makes agentic benchmark numbers hard to compare across reports. A TerminalBench-Lite figure is now underdetermined unless the action-selection policy is stated.
The distillation result is the practically interesting one: it converts a dependency on a frontier verifier into weights in a 9B model. But see Open Questions — the headline configuration still requires calling GPT-5.6 Sol on every step.
Open Questions
- The 68.03% configuration uses a frontier model as verifier on every action. No cost accounting for that is given — the token-cost comparison reported is between action scaling and trajectory scaling, not between this and the base agent.
- TerminalBench-Lite is the only benchmark named. "Additional models, benchmarks, and harnesses" is asserted without specifics, so the breadth claim cannot be checked.
- TMAX-9B is not a model this wiki holds a page for, and nothing read establishes who makes it or what it is. The generator's identity bounds how much the 50.00% baseline means.
- How action verification interacts with irreversible actions is not addressed — verification before execution is most valuable exactly where a mistake cannot be undone, and no such split is reported.
Cite
arXiv:2609.39982, published 2026-09-30. Surfaced via HuggingFace Daily Papers, 2026-10-02, 98 upvotes — a popularity signal from that community and nothing more (source).
Captured 2026-10-04, +4 days, from the owed 2026-10-02 snapshot.