$ cat briefs/daily/2026-10-07.md
2026-10-07
October 7, 2026 (Wed)
3 stories · 2 paper picks · 4 new pages
Top Stories
1. Ironclad makes workflow requirements the training target
OpenAI reports Astra scoring 55.0% versus GPT-5.6 Sol's 41.6% on 11 contracting tasks. These are rubric scores; reasoning settings differ. Reported 19.2 versus 37.0 minutes per attempt are simulated estimates, not measured customer savings.
Why it matters: software partners can supply both realistic practice environments and explicit success criteria. → Astra (source)
2. Anthropic expands verified cyber access
CVP now has Defense, Red Team and Specialized Access tiers. On CyScenarioBench, Opus 5.5's Defense tier blocks 46 of 50 trials; Red Team blocks none and completes 34 of 50.
Why it matters: safeguard settings materially change the observed result for the same model. → AI-Enabled Cyberattacks (source)
3. Mistral Large 4 enters public preview
The API is available; weights are promised by the end of October. Mistral reports 61.7% DeepSWE v1.1 and 59.9% AutomationBench. Its launch and documentation disagree on parameter counts; the wiki preserves both.
Why it matters: this adds a European option for agent workloads, while downloadable weights remain a future commitment. → Mistral Large 4 (launch) (documentation)
Paper Picks
UndoBench — completion does not establish safe recovery. Nominal competence reaches 83.54%, conditional recovery 46.72%, and naive retries duplicate external effects in 53.33% of trials. Read it for counterfactual fault testing; results depend on when mutation occurs. → UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents (source)
Self-generated feedback can destabilize test-time training. Freezing the training-text generator removes over 98% of damage in two tested configurations. Read it for causal controls and checking independent real text before retaining weight updates; this is not a finding against every form of extra inference compute. → Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation (source)
Watch
- EmbeddingGemma 2: Google releases a 740M, Apache 2.0 multimodal embedder with modular encoders and 8K context. A useful local-retrieval option; reported memory depends on the device and quantization. → EmbeddingGemma 2 (source)
- Mathematics disclosures: OpenAI announces Lean formalizations, reasoning summaries and compute estimates for new internal-model results. The announcement does not establish that every proof is formalized. → AI for Mathematics (source)
New in Wiki
- Mistral Large 4 and EmbeddingGemma 2 — release records and source limits. (Mistral) (Google)
- UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents and Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation — methods and unanswered questions. (recovery) (adaptation)
Updates
- Agents (LLM Agents) connects workflow scoring with recovery evaluation. (source)
- Test-Time Compute (Inference-Time Compute Scaling) distinguishes fixed-weight inference from weight-changing adaptation. (source)