$ cat briefs/daily/2026-10-10.md
2026-10-10
October 10, 2026 (Sat)
2 stories · 2 paper picks · 4 new pages
Top Stories
1. Asana's browser-agent savings start with keeping the cache reusable
Asana's 144-run study combines history caching, batched screenshot removal and a larger text budget. Optimization alone cuts cost 29x on the original model; the headline 76x also switches to GPT-6.1 Sol. Capped baseline runs make the reductions lower bounds, and small sample sizes limit fine comparisons.
Why it matters: history policy can dominate agent economics, so cost comparisons must name both model and workflow. → Asana (source)
2. Qwen releases an eight-step image checkpoint and hosted APIs
Qwen-Image-2.1-Turbo generates and edits images with 8 denoising steps; Pro/Turbo APIs are announced as live. The repository's research license is non-commercial, but the separate Turbo model card was not inspected; checkpoint-specific terms and API pricing remain unverified.
Why it matters: downloadable and hosted deployment paths are available, while step count alone establishes neither latency nor quality. → Qwen-Image-2.1-Turbo (release) (repository license)
Paper Picks
TestPrism — a test must accept other correct implementations. Across fourteen baseline configurations, joint success is 28.00%, versus 59.67% with one reference. Read it for evaluating both missed bugs and unjustified failures; TestHelix's reported improvement lacks a cost-matched comparison in the abstract. → TestPrism: Rethinking Test Evaluation Beyond a Single Reference (source)
Memento 3 — revise the rulebook while keeping the LLM fixed. Verified executable world models clear all levels of 25 public ARC-AGI-3 games, with mean RHAE 100.0 using 44% of human actions. Read it for replay-based verification; this is not a private-test result or a total-cost measurement. → Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks (source)
Watch
- REMORY: learned soft memory supplements a textual summary. On SummHay, the authors approach the full-context joint score using 5.2% of input positions; training and end-to-end cost remain unspecified in the abstract. → Context Compaction (source)
New in Wiki
- Asana — new organization page for reader review. (source)
- Qwen-Image-2.1-Turbo — release record with explicit unknowns. (source)
- TestPrism: Rethinking Test Evaluation Beyond a Single Reference and Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks — methods, results and open questions. (tests) (rulebooks)
Updates
- Astra records experiment execution with human review and retained traces. (source)
- World Models distinguishes replay consistency from unseen-state generalization. (source)