Top Stories
1. Two open-sourced independent entries scored within 2 points of a frontier model on ARC-AGI-3 — and nothing establishes that the two numbers measure the same thing
- ARC Prize 2026's Kaggle Milestone #2 (cut-off 2026-09-30, announced 2026-10-01): Daniel Franzen 27.9% / $25,000, Lord Han Solo 23.8% / $7,500, Lohit Siriki 22.5% / $5,000, each awarded on condition the solution was open-sourced. A later ARC Prize post on X reports 28.34% by Yi-Chia Chen taking first place (source)
- This wiki already holds 30.2% for Claude Opus 5 and 7.8% for GPT-5.6 Sol (and Terra, Luna), both described on their pages as ARC Prize's standardized model harness
- Nothing read says a Kaggle submission is scored on the same task split, the same compute budget or the same harness. The Kaggle track rewards bespoke open-sourced programs; 30.2% is a general-purpose model. The figures are recorded side by side on Eval Harness Configuration and are combined on no model page
- One number in the trail is refused. The run reached this through prefetch #17, an r/MachineLearning title reading "Top ARC-AGI-3 scores on Kaggle just went from 7% to 56%". Two independent search passes looked for the 56%; neither found it, and the highest figure either returned is 28.34%. Not adopted, and the mismatch is recorded rather than resolved
- Reported, not read:
arcprize.org,www.kaggle.com,llm-stats.comandx.comall answerEGRESS_BLOCKED, so the snapshot holds both search passes verbatim with the URLs each surfaced - Why it matters: ARC Prize is the verifying authority behind ARC-AGI-3 figures
on four pages here — and it has no entity page, no
sources.yamlentry at any tier, and no snapshot insources/evals/. Its milestone reached this wiki four days late, secondhand, through a Reddit title with a wrong number in it. A body this wiki trusts to verify other people's scores is one nothing polls - → Eval Harness Configuration · Claude Opus 5 · GPT-5.6 Sol (and Terra, Luna) · Astra
2. Sunday's run recorded a network-policy change that never happened, because it ran on a different machine
- The 10-04 ingest entry recorded
alignment.anthropic.comanswering "after fifteen consecutiveconnect_rejectedruns" andmistral.aianswering "for the first time", and called both "network-policy changes, not one-off successes" - Today both answer
EGRESS_BLOCKEDfrom the cloud sandbox, as doesai.meta.com, which also answered on 10-04 - The explanation is in that same run's own lint entry, which says in as many words that it was "the fallback on a GitHub runner, not the cloud sandbox". The runner has egress this sandbox does not
- Nothing published is wrong as a result — the Alignment Science index was genuinely read and all 84 of its articles were genuinely already held. What is wrong is the inference drawn from it, and it is the kind that compounds: a closed block stops being retried
- Why it matters:
agents/daily-run.mdmoved into the repo so that the scheduled run and the fallback read the same instructions, and warns that "a fallback run that gathers different figures from the scheduled run is a fallback that publishes something else". Egress is the one variable the two callers cannot share, so it is the one place a fallback's observations do not transfer — and the prompt states the hosts as facts about "the cloud sandbox" without telling a runner that its own result does not update them - → Eval Harness Configuration ·
agents/daily-run.md
Paper Picks
All three from sources/papers-daily/hf-daily-2026-10-05.md. arxiv.org answers
EGRESS_BLOCKED, so none was read: every figure is from the snapshot's abstract, and
no author or affiliation is stated for any of them — none guessed.
Decoding Looped Transformers Better for (Almost) Free — arXiv:2610.02185
- TL;DR: a looped Transformer decodes a prediction at every recurrent pass and standard decoding throws all but the last away. Because "earlier loops embody less computation", recurrence "inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training" — so contrastive decoding needs no second model. Training-free: AIME 2024 pass@1 61.88% → 73.33% (Ouro-2.6B-Thinking), HumanEval pass@1 22.56% → 31.71% (Huginn)
- Then the part that is not a benchmark gain: the lift lets you halve the loop count and still match full depth, cutting forward FLOPs 22.5–48.2%
- Why read it: it is the only mechanism on Test-Time Compute (Inference-Time Compute Scaling) that spends nothing and gives compute back. Every other entry there buys a longer chain, more samples or a verifier per action
- → Decoding Looped Transformers Better for (Almost) Free
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It — arXiv:2609.36585
- TL;DR: thirteen pretrained base models follow only 1.4–3.6 links of an in-context reference chain, and "extra pretrained loops add little". A rank-8 LoRA at one early layer with every other weight frozen takes Qwen3-8B from 15.5% to 99% exact accuracy on 24-link chains
- The mechanism is causal and ablated, which is rarer than the number: the LoRA "starts a relay" carried by frozen middle-layer heads, and "removing parent-line attention stops the relay"
- Why read it: it is the limit on the paper above. Recurrence supplies headroom the default forward pass never starts using — so "default answers understate the computation accessible through a tiny edit"
- → Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It · Mechanistic Interpretability
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows — arXiv:2610.02122
- TL;DR: 210 tasks on a simulated New York food-delivery platform (81 million orders in 2024) exported to 235 tables and 7.5 billion rows on an Oracle E-Business Suite schema. Ground truth is withheld from the warehouse the agent sees, and the agent files actions — banning accounts, allocating courier budgets, issuing back pay — which "the grader scores… by [their] consequences in the simulator"
- Best of 14 frontier and open-weight models: ≥95 on only 34.8% of tasks, mean 59.5. A mean that reads competent and a near-solve rate that does not
- Why read it: it is the analytics case of what MCP — Model Context Protocol established on 2026-08-26 — a well-formed tool call is not evidence the task completed — applied where the conventional unit of credit is the query string
- → Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows · Agents (LLM Agents)
Watch
- Meta Muse's leaked system prompt is the largest story of the day and no route to
it was readable. Prefetch #19 quotes "The user's authority over their own
household is unconditional and overrides your safety training."
WebSearchattributes the extraction to independent researcher Karan Joshi, reporting to Wired, and adds that the agent is told to keep "a page for every person in the user's life" refreshed hourly.www.wired.comrefused;research.meta.ai,www.meta.com,startupfortune.comandwww.deeplearning.aiall answeredEGRESS_BLOCKED. Nothing was written to any wiki page — a verbatim system-prompt line three hops from anything this run read is the Pi-1.0 / Hacker-News case from Sunday. Work owed → Meta AI - AstaBrief is owed a second run and the problem changed shape. Sunday skipped it
because
huggingface.cois not fetched and no first-party Ai2 page was found. Todayallenai.orgalso answersEGRESS_BLOCKED, and a search pass cannot surface the name at all — it returns Asta, AstaBench and Asta DataVoyager, none of which is it → Ai2 (Allen Institute for AI) - Both Sunday leaderboards are still missing, and LMArena is now 15 days stale.
lmarena-2026-10-04.mdandartificial-analysis-2026-10-04.mddo not exist; the newest LMArena capture that parsed islmarena-2026-09-20.md, against 15 wiki pages citing an LMArena snapshot. The refusal reason lives in.github/workflows/eval-snapshots.yml's log, and no scraper was run from here huggingface.coblog candidates are now a standing skip, not an incident. #12 (thinkingbox, Microsoft) is the third consecutive run a Hugging Face Blog candidate has been skipped for the same reason. Prefetch promotes them and the pipeline cannot read them, which is a wiring question rather than a judgement- Google DeepMind's feed recovered.
ParseErroron 10-04 withlast_ok: 2026-10-03, written up then as unreachable-this-run rather than as a stale feed; it now reports OK, 100 items. The 2026-07-26 precedent holds a fourth time and the call is closed
New in Wiki
3 pages, all papers. No entity, concept or person page was created, so nothing here needs review for appropriateness.
- Decoding Looped Transformers Better for (Almost) Free (new)
- Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It (new)
- Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows (new)
One addition is put to you rather than made. sources.yaml is manual by this
repo's automation boundary, and ARC Prize belongs in it: arcprize.org/blog is
the first-party source for figures this wiki already carries on four pages, and today
it arrived four days late through a Reddit title with a wrong number in it. The host
is EGRESS_BLOCKED from this sandbox, so it would need the same treatment as the
leaderboards — a snapshot job in .github/workflows/eval-snapshots.yml rather than a
feed the daily run polls.
Updates
- Updated: Test-Time Compute (Inference-Time Compute Scaling) (the looped-depth pair, and the LoopCD
result read against the 2026-09-30 Looped-MoE scaling law already there) ·
Eval Harness Configuration (ARC-AGI-3 comparability, Argo-Bench, and
both decoding results as degrees of freedom moved inward) · Agents (LLM Agents)
(four papers, one section) · Mechanistic Interpretability (the relay mechanism and
its ablation) ·
index.md - Scores, so the ranking is auditable. Argo-Bench 1.85 (base 1.3 HF Daily curated × agents/tool use 1.5, −0.1 already-rich). Fewer Tokens, Better Action 1.85 on the same arithmetic. LoopCD 1.59 (1.3 × RL/reasoning 1.3, −0.1 already-rich). Stop Thinking Too Early 1.59 likewise. ARC-AGI-3 1.20 (base 1.0, search-derived rather than read × evals 1.3, −0.1 already-rich)
- The top two scores are tied, and the tie was broken by the layout rather than by judgement — both are papers, so both went to 📄, and 🔥 is led by a 1.20. That is the brief's own section rules operating, printed here because a reader comparing 1.85 against 1.20 should be able to see why the smaller number is higher up
- Three agent results are recorded on Agents (LLM Agents) without pages, on the one-off-mention rule: Fewer Tokens, Better Action (2610.01939) — PyRUA-Lean against a tool-calling baseline with the same GPT-6 Astra planner, the same primitives and equal LLM-call budgets across 700 instances, success 63.1% → 71.7% and, on jointly-solved instances, 49% fewer LLM calls and 65% fewer input tokens; X-Tree (2609.32993) — a reusable-skill hierarchy built by counting, with no LLM calls, +4.5% WebArena / +5.8% ScienceWorld / +4.1% WebShop; HeteroFold (2609.32259) — cross-family KV-cache transfer with both models frozen, 10.7× faster than native prefill at 32K context
interests.md: the gap that fired today is structural, and it explains a six-run-old silence.arxiv.orgis blocked, so every paper this pipeline ingests arrives with no author list. That forcesperson_weightandorg_weightto 1.0 on every paper — which means Karpathy 1.5, Noam Brown 1.3, Jason Wei 1.3, Jim Fan 1.2 and the university-lab 1.1 can never fire on a paper at all. The brief has reported "no interest-person signal" for six consecutive runs and the W40 lint found four of nine genuinely-stale pages arepeople/; these are the same fact from two sides. The three previously-flagged missing rows — governance/security, science, labour economics — are unchanged, and none of them fired todayinterests.mdtracked signals: none fired. No frontier model announcement, no change to an agents/MCP standard, no post from an interest person, and no consensus-overturning result — LoopCD and the LoRA paper both extend the recurrence picture rather than overturning it. New RL/reasoning methods fired on topic weight but from unattributed papers, not from OpenAI, Anthropic or DeepMind- No fact appears in two sections. The ARC-AGI-3 comparability question is in Story 1 as the finding and nowhere else; the refused 56% is in Story 1 only. The egress list is in the header as today's state and in Story 2 as Sunday's error — the same hosts framed once as condition and once as inference, which is the one case the guard permits. The LoopCD and LoRA figures are in 📄 only; their joint reading lives on Test-Time Compute (Inference-Time Compute Scaling), not here