$ cat briefs/daily/2026-10-08.md
2026-10-08
October 8, 2026 (Thu)
3 stories · 2 paper picks · 5 new pages
Top Stories
1. Haiku 5.5 lowers the entry price for bounded agent work
Anthropic charges $0.10/M input and $0.50/M output for prompts up to 100K tokens; above that, rates rise to $0.50/$2.50. Its launch reports 72.4% on OSWorld 2.1's offline subset. Migration needs care: the same text uses approximately 30% more tokens than Haiku 4.5, and thinking configuration changes.
Why it matters: cheaper subagents become practical, but short-prompt prices and benchmark subsets must stay attached to the claims. → Claude Haiku 5.5 (launch) (migration)
2. Liquid AI releases open decision models
d1-3B handles text/images; experimental d1-omni-600M handles text with images or audio. Both return decisions without generating tokens. Liquid reports 48.57 and 15.95 on the Decision Index v0.2.1 public split; no private vision-split result is published.
Why it matters: local systems can make structured multimodal decisions without a text-generation loop. → d1-3B · d1-omni-600M (source)
3. GPT-6 brings generated interfaces into ChatGPT Chat
Intelligent UI combines text with interactive components rendered progressively. Paid-tier rollout began October 7; Free/Go starts October 8, using conversation-tuned Sol and Luna respectively. Work and Codex models are unchanged.
Why it matters: answers can become usable interfaces within the conversation; this announcement changes Chat availability rather than the API rate card. → OpenAI (source)
Paper Picks
CheckerBench — build the checker, not just the patch. Across 300 tasks and 21 model-harness configurations, mean Pass@1 is 32.30%, best 45.33%. Read it for independent rebuilding and false-positive evaluation; the abstract does not identify the winning configuration. → CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers? (source)
SafeActBench — did evidence precede action? 656 cases track investigation, prerequisites and execution. Static judgment can look strong while interactive action fails. Read it for the distinction between a correct outcome and a justified action; the abstract gives no numerical success rate. → From Evidence to Action: How Tool-Using Agents Fail (source)
Watch
- TRACE: rollout-guided FP4 training reports up to 5.4x rollout speedup with RL performance comparable to BF16 across four MoE models. This is not an end-to-end training speedup. → Agentic Reinforcement Learning (source)
New in Wiki
- Release records: Claude Haiku 5.5, d1-3B, d1-omni-600M. (Anthropic) (Liquid AI)
- Research notes: CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?, From Evidence to Action: How Tool-Using Agents Fail. (checker synthesis) (evidence and action)
Updates
- Claude Sonnet 5.5 cache reads fell from $0.20/M to $0.10/M. Anthropic estimates around 20% lower cost on most agentic tasks; savings depend on the workload. (source)