$ tree wiki/
Wiki
The compounding AI-trend knowledge base: organizations, models, concepts, people, and papers in one place.
409 pages
Entities
/35AI Evaluator Forum (AEF)A consortium of AI research organisations focused on independent, third-partyupdated 2026-09-208 in-linksAi2 (Allen Institute for AI)Ai2, the Allen Institute for AI, is a research institute that publishes openupdated 2026-10-021 in-linksAleph AlphaAleph Alpha is a German AI lab. This page was created on 2026-10-04 on theupdated 2026-10-043 in-linksAlibaba / Qwen AI LabHangzhou-based Chinese technology conglomerate. In AI: principally known for the Qwen family of open-weight l…updated 2026-10-0143 in-linksAMDAdvanced Micro Devices — the second-source supplier of AI datacentre accelerators,updated 2026-09-304 in-linksAnt Group (inclusionAI / AntLing)Ant Group publishes open-weight models through inclusionAI, its open-sourceupdated 2026-09-276 in-linksAnthropicSan Francisco-based AI safety company. Develops the frontier Claude model series — models/claude-opus-5, mode…updated 2026-10-0491 in-linksApodexA lab building what it calls "Discoverative AI" — AI that makes newupdated 2026-08-264 in-linksAppleConsumer electronics and software company. Not a frontier AI lab, but as of WWDC 2026 (June 8), Apple has rep…updated 2026-07-125 in-linksDeepSeekChinese AI research lab (affiliated with High-Flyer Capital Management, Hangzhou). Known for releasing fronti…updated 2026-10-0119 in-linksDiscovery LoopA Public Benefit Corporation founded 2026-08-05 by four departing Googleupdated 2026-08-073 in-linksGoogle DeepMindGoogle's AI research division (the 2023 merger of DeepMind + Google AI). Develops the Gemini model series. Fo…updated 2026-10-0276 in-linksHugging FaceThe Hub — the repository where open-weight models, datasets and demo Spacesupdated 2026-09-2512 in-linksIBMEnterprise technology company shipping the Granite family of open-weightupdated 2026-08-274 in-linksInstitute of Foundation Models (IFM)A research lab at MBZUAI (Mohamed bin Zayed University of Artificialupdated 2026-09-044 in-linksLiquid AILiquid AI publishes the LFM (Liquid Foundation Model) series — small open-weightupdated 2026-08-214 in-linksMeituanChinese consumer technology conglomerate (food delivery, local services, travel). In 2026, Meituan's internal…updated 2026-07-054 in-linksMeta AIMeta's AI research division. FAIR (Fundamental AI Research) plus the recently established Meta Superintellige…updated 2026-10-0128 in-linksMicrosoftTech giant based in Redmond, WA. Dominates the AI developer and enterprise tooling market through GitHub Copi…updated 2026-09-1517 in-linksMiniMaxShanghai-based AI startup (founded 2021). Develops frontier multimodal models and consumer AI products. Best …updated 2026-09-2910 in-linksMistral AIParis-based AI research and product company. Open-weight model lineup (Mistral, Mixtral) + Le Chat chatbot + …updated 2026-08-2920 in-linksMoonshot AIBeijing-based AI startup, creators of the Kimi assistant and model family. Known for pushing long-context cap…updated 2026-10-0121 in-linksNVIDIAA GPU manufacturer and the de facto standard for AI compute infrastructure. The infrastructure core of the AI…updated 2026-09-3031 in-linksOpenAISan Francisco-based frontier AI lab. Developer of the GPT series (ChatGPT). 2026 slogan: "Year of Science". F…updated 2026-10-0179 in-linksPoolsideAI lab building models for agentic coding. Released models/laguna-s-2-1 onupdated 2026-08-225 in-linksPrismMLA compression lab, described in one source as a Caltech spin-out, that publishes ternary-weight versions of o…updated 2026-09-197 in-linksRunwayRunway is a research company building generative video and audio models,updated 2026-09-268 in-linksSakana AIA Tokyo-based AI R&D company, founded 2023 by former Google researchers,updated 2026-09-127 in-linksTencentTencent is a Chinese consumer-internet and cloud company whose AI research runsupdated 2026-09-279 in-linksThinking Machines LabAI lab whose first production model, models/inkling, shipped on 2026-07-15 as anupdated 2026-08-033 in-linksTypeSafe AIA San Francisco company that emerged from stealth on 2026-09-15 withupdated 2026-09-244 in-linksWorld LabsSpatial-intelligence models that generate, reconstruct and simulate interactiveupdated 2026-09-303 in-linksxAIAI research company founded by Elon Musk (2023). Develops the Grok model series. Leverages real-time data acc…updated 2026-10-0125 in-linksXiaomiChinese consumer-electronics manufacturer whose MiMo model family reachedupdated 2026-09-236 in-linksZ.aiZ.ai is the international brand of Zhipu AI (智谱AI), a Beijing-based AI company founded in 2019, spun out of T…updated 2026-10-0124 in-links
Models
/121AlphaEvolveDesigning advanced algorithms via evolutionary searchupdated 2026-07-2112 in-linksAstraOpenAI's next major model, named publicly for the first time on 2026-08-01 in aupdated 2026-09-3039 in-linksChatGPT Images 2.5Pricing is unknown on purpose and it is not the same as free. The two APIupdated 2026-09-094 in-linksClaude Fable 5June 9, 2026 — the first publicly available Mythos-class model from Anthropic. Fable 5 and Claude Mythos 5 sh…updated 2026-09-0239 in-linksClaude Fable 5.1Anthropic's 2026-09-01 point release of the Fable line, shipped together withupdated 2026-09-1315 in-linksClaude Mythos PreviewAnnounced April 7, 2026 alongside Project Glasswing. Withheld from commercial release. Anthropic explicitly s…updated 2026-10-0119 in-linksClaude Opus 4.7Thoroughness / consistencyupdated 2026-08-1412 in-linksClaude Opus 4.82026-05-28 — released alongside the Anthropic $65B Series H funding announcement.updated 2026-07-2722 in-linksClaude Opus 5Max output (sync): 128k tokensupdated 2026-08-2443 in-linksClaude Opus 5.5Anthropic's 2026-09-22 release of the Opus line, and the first model on thisupdated 2026-09-2922 in-linksClaude ScienceClaude Science is a scientific research workbench that gives researchers a unified environment for computatio…updated 2026-07-213 in-linksClaude Sonnet 5June 30, 2026 — released the same day as Fable 5/Mythos 5 export-control restoration and Claude Science launc…updated 2026-09-2919 in-linksClaude Sonnet 5.5Anthropic's 2026-09-28 Sonnet release, six days afterupdated 2026-09-308 in-linksCo-Scientist (Google DeepMind)Research demo / early access: 2026-02 (initial Co-Scientist blog)updated 2026-07-216 in-linksCosmos 3 Super2026-06-01 (Cosmos 3 series launched together: Super / Nano / Edge pending)updated 2026-07-215 in-linksCosmos-H-DreamsCosmos-H-Dreams is an action-conditioned world foundation model for surgical robotics simulation. It generate…updated 2026-07-287 in-linksDeep Research MaxAnnounced April 22, 2026 alongside Deep Research (the speed-optimized tier) and the Gemini Enterprise Agent P…updated 2026-07-212 in-linksDeepSeek V4Preview: April 24, 2026. GA: mid-July 2026.updated 2026-08-148 in-linksDeepSeek V4-Flash2026-07-31, moving V4-Flash out of Preview into public beta under the buildupdated 2026-08-2215 in-linksDeepSeek V4-Flash-Vision-ExpEvery unknown above is genuinely unstated in the coverage read, not omitted.updated 2026-08-223 in-linksDeepSeek V4-Pro-0813This page published a retirement that is not happening, and the correction isupdated 2026-09-1315 in-linksDeepSeek V4.1-FlashTwo prefetch candidates on the 2026-09-10 run named this model a day before itupdated 2026-09-1910 in-linksDevstral 2May 2026 (exact date TBD — flagged in sources as a May 2026 release)updated 2026-07-215 in-linksDiffusionGemmaGoogle DeepMind's first open-weight text diffusion model, built on the Gemma 4 26B MoE backbone.updated 2026-07-273 in-linksFugu MaxLicense reads unknown rather than proprietary because nothing read names aupdated 2026-09-127 in-linksFugu Ultra v2License reads unknown because nothing read names a license for any Fuguupdated 2026-09-125 in-linksGemini 3.1 Deep ThinkAdvanced "Deep Think" version: ~January 2026updated 2026-07-219 in-linksGemini 3.5 FlashOutperforms Gemini 3.1 Pro across the full coding and agentic benchmark suite.updated 2026-09-137 in-linksGemini 3.5 Flash CyberChrome V8 JavaScript engine vulnerability finding:updated 2026-09-038 in-linksGemini 3.5 Flash-LiteHigh-throughput, low-latency agentic pipelines (agentic search, document processing)updated 2026-09-135 in-linksGemini 3.5 ProAnnounced at Google I/O 2026 (May 19, 2026) with a June 2026 general-availability target. As of July 17, 2026…updated 2026-07-2213 in-linksGemini 3.5 TranscribeContext window is unknown because nothing read states an audio-length or tokenupdated 2026-08-276 in-linksGemini 3.6 FlashKnowledge cutoff: March 2026 (up from January 2025 on 3.5 Flash).updated 2026-09-1310 in-linksGemini 3.7 FlashGenerally available at announcement — a stable API model, not a previewupdated 2026-09-1311 in-linksGemini 3.8 FlashGenerally available at announcement, API model id gemini-3.8-flashupdated 2026-09-1313 in-linksGemini 3.8 Flash CyberContext window is unknown rather than 1,048,576. Its general-purposeupdated 2026-09-035 in-linksGemini 3.8 Flash TTSentities/google-deepmind's expressive text-to-speech model, announcedupdated 2026-09-243 in-linksGemini 3.8 Flash-Lite TTSentities/google-deepmind's high-volume text-to-speech model, announcedupdated 2026-09-243 in-linksGemini 3.8 LiveAnnounced together with models/gemini-3-8-live-extended-thinking in a singleupdated 2026-09-258 in-linksGemini 3.8 Live Extended ThinkingAnnounced together with models/gemini-3-8-live in a single post,updated 2026-09-162 in-linksGemini 42026-09-30 — the generation's first named model shipped asupdated 2026-10-025 in-linksGemini 4 Argonentities/google-deepmind's 2026-09-30 frontier release, announced by SVPupdated 2026-10-025 in-linksGemini OmniThis page is the I/O announcement and stays one. The Omni line has sinceupdated 2026-08-286 in-linksGemini Omni 1.1 FlashAnnounced and available the same day, described as a production-ready update forupdated 2026-08-284 in-linksGemini Robotics 2Announced 2026-07-28 as the vision-language-action member of a three-model family. Theupdated 2026-07-317 in-linksGemini Robotics ER 1.6Released April 15, 2026. Successor to Gemini Robotics-ER 1.5. Notable collaboration with Boston Dynamics on i…updated 2026-07-218 in-linksGemini Robotics ER 2Announced 2026-07-30, two days after its VLA siblingupdated 2026-07-317 in-linksGemini SparkGemini Spark is a 24/7 cloud-based personal agent that takes actions on behalf of users even when they're off…updated 2026-07-212 in-linksGemma 3nEarly preview released 2026-05-12.updated 2026-07-213 in-linksGemma 4 12BFull precision: 16GB VRAM (RTX 4060, RTX 5090, Apple Silicon Mac)updated 2026-07-218 in-linksGLM-5.2June 13, 2026: Available to Z.ai GLM Coding Plan subscribersupdated 2026-09-0620 in-linksGLM-5.3Context window moved twice, and the second move is the one that settles it.updated 2026-10-0129 in-linksGLM-5.3-Flash320B total / 18B active MoE (320B-A18B), natively multimodal — image andupdated 2026-09-2911 in-linksGPT-5.5 InstantA claim of simultaneous improvement across three axes (intelligence, clarity, personalization).updated 2026-08-094 in-linksGPT-5.6 Sol (and Terra, Luna)GPT-5.6 is OpenAI's three-tier model family, announced June 26, 2026. Terra and Luna wereupdated 2026-09-1343 in-linksGPT-5.6-CyberBuilt on top of Sol, trained to improve at findingupdated 2026-08-129 in-linksGPT-6 LunaOpenAI's 2026-09-22 fast, cost-efficient model of the GPT-6 series, releasedupdated 2026-09-2310 in-linksGPT-6 SolOpenAI's 2026-09-22 cost-efficient high-end model of the GPT-6 series,updated 2026-09-3016 in-linksGPT-6.1 SolOpenAI's 2026-09-29 DevDay release, seven days afterupdated 2026-09-308 in-linksGPT-Image-2.5 FlareThe Pricing row is the published API rate carried by two independent passes,updated 2026-09-093 in-linksGPT-Image-2.5 SunburstThe Catalogue id row is present because the model's API page path names itupdated 2026-09-093 in-linksGPT-Live-1Full-duplex speech: model listens and speaks at the same time; users can interrupt naturally mid-sentenceupdated 2026-07-217 in-linksGPT-Realtime-2 (OpenAI)2026-05-07. Simultaneously: Realtime API exits beta → generally available for production.updated 2026-07-215 in-linksGPT-Rosalind2026-04-16 (model release); 2026-05-29 (Biodefense program launch)updated 2026-07-215 in-linksGranite 4.2Three variants in one release: 3B for edge devices, 8B mid-range, 30Bupdated 2026-08-272 in-linksGrok 4.1 Fast (xAI)May 2026 (exact date unconfirmed; live on x.ai/api as of May 2026)updated 2026-07-213 in-linksGrok 4.5xAI's enterprise deployment of the V9-class model. 1.5 trillion parameters, Cursor-trained — entered private …updated 2026-07-275 in-linksGrok 4.6Three rows moved off unknown on 2026-08-16, and none of them from aupdated 2026-08-2316 in-linksGrok 4.7entities/xai's flagship coding and knowledge-work model, releasedupdated 2026-09-256 in-linksGrok Build2026-06-22: /goal mode — long-running autonomous execution (plan→execute→verify) for SuperGrok/X Premium+updated 2026-07-277 in-linksGrok Imagine Image 2.0Pricing and License are unknown because no first-party page was readableupdated 2026-08-106 in-linksGrok Imagine Video 1.5 (Preview)2026-06-03 (API preview)updated 2026-07-215 in-linksGrok V9-MediumxAI's coding-focused foundation model — 1.5-trillion parameters, ~3× larger than the prior production Grok. T…updated 2026-07-214 in-linksGrok Voice Think Fast 2.0Neither a parameter count nor an architecture nor a context window was published in anyupdated 2026-08-056 in-linksGrok Voice Transcribe 2.0Context window reads unknown and that is the finding. No maximum audioupdated 2026-09-253 in-linksGWM Worlds 2Four rows need their reading statedupdated 2026-09-2611 in-linksHy4 previewContext window is recorded as the coverage states it — "exceeding 1Mupdated 2026-09-138 in-linksInklingPricing is unknown because Thinking Machines publishes no list price that wasupdated 2026-08-0311 in-linksJevThree of these rows need their reading stated, because each is unknown or oddupdated 2026-09-245 in-linksK2 HorizonA family of six models from the Institute of Foundation Models, releasedupdated 2026-09-045 in-linksKimi K2.8 PreviewPricing is unknown and the gap is documented rather than assumed: as ofupdated 2026-09-205 in-linksKimi K3Frontend Code Arena: beats Anthropic Fable 5 (human-preference Elo) — Moonshot's reported benchmarkupdated 2026-09-3035 in-linksKolibri-1entities/aleph-alpha's 2026-10-03 open-weight release, and the firstupdated 2026-10-044 in-linksLaguna S 2.1Pricing is unknown: Poolside published weights, not an endpoint price, and theupdated 2026-08-0310 in-linksLeanstral 1.5Formal verification: generating Lean 4 proofs for functions and algorithmsupdated 2026-07-217 in-linksLFM2.5-2.6BTwo rows need their unknown explained, because neither is an unread field:updated 2026-08-217 in-linksLing-3.0-tinyThree rows need their reading stated:updated 2026-09-278 in-linksLongCat-2.0LongCat-2.0 had been running quietly on OpenRouter under the codename "Owl Alpha" before its identity was rev…updated 2026-07-285 in-linksLyria 3.5Announced and rolled out on the same day; no separate preview period was statedupdated 2026-07-301 in-linksMAI-Code-1 / MAI-Code-1-FlashMAI-Code-1-Flash: 2026-06-02 (immediately available)updated 2026-07-217 in-linksMAI-Thinking-12026-06-02 (announced at Build 2026). Exact API/GA date not announced.updated 2026-07-216 in-linksMiMo-V2.6-Proentities/xiaomi's 2026-09-22 flagship: a 1.02T-parameter sparseupdated 2026-09-305 in-linksMing-Image-0.1-DesignFour rows need their reading statedupdated 2026-09-264 in-linksMiniMax H3Context window is unknown: this is a video generation model billed per outputupdated 2026-08-047 in-linksMiniMax M3June 1, 2026 — MiniMax official blog and HuggingFace release. Captured by this wiki July 17 due to WAIC 2026 …updated 2026-07-286 in-linksMiniMax Music 3.0Context window is unknown: this is a text-to-music model whose inputs areupdated 2026-08-182 in-linksMistral Large 3Early access opened July 6, 2026. CEO Arthur Mensch confirmed the model on July 4, 2026 in a TechCrunch profi…updated 2026-07-215 in-linksMistral Medium 3.5Pending further detail from the full announcement materials.updated 2026-07-284 in-linksMuse Glimmer30B dense multimodal, built to run offline on consumer hardware. 4-bitupdated 2026-08-1110 in-linksMuse ImageArena text-to-image: #2 (human-preference Elo at launch)updated 2026-07-204 in-linksMuse Realtime AvatarFour rows need their reading statedupdated 2026-09-275 in-linksMuse Spark (1.0 / 1.1)Major upgrade released alongside the Meta Model API — Meta's first paid external AI product.updated 2026-08-079 in-linksMuse Spark 1.2There are two price tiers, and the cheaper one is paid for in training data. Theupdated 2026-08-106 in-linksMuse Spark 1.3Input is text, image and video; output is textupdated 2026-09-045 in-linksMuse VideoPreviewed July 7, 2026. General availability date not announced.updated 2026-07-204 in-linksMuse Voice TranscribeInput is streaming audio; output is text with speaker labels andupdated 2026-09-073 in-linksNano Banana 2 Lite (Gemini 3.1 Flash Lite Image)June 30, 2026 — the fastest and most cost-efficient model in the Nano Banana (Gemini image generation) family…updated 2026-07-213 in-linksNemotron 3.5 Lightning30B total / 3B active hybrid Mixture-of-Experts, described as interleavedupdated 2026-08-1211 in-linksNVIDIA Kumo TabularAn open foundation model for tabular classification and regression, publishedupdated 2026-09-303 in-linksProject PolarisPre-announcement: 2026-06-01 (Microsoft Build 2026, June 2-3 keynote)updated 2026-07-216 in-linksQwen 3.8 27BIt shipped. Released 2026-08-14 at 15:00 UTC — 27.78B denseupdated 2026-08-2323 in-linksQwen 3.8 MaxFour rows on this page read unknown from 2026-07-20 until 2026-08-03. They wereupdated 2026-09-0318 in-linksQwen-Drive-1.0-4B2026-09-07, four days before this captureupdated 2026-09-117 in-linksQwen-Image-2.1Context window is unknown and the gap is a category mismatch rather than aupdated 2026-09-216 in-linksQwen3.8-Flash-NextThe release landed on the date the countdown named. Yesterday this page readupdated 2026-09-029 in-linksRobostral Navigate2026-07-08. Announced via Mistral blog (source) and covered by Bloomberg.updated 2026-07-215 in-linksShieldstral 1.0Two unknown rows, both genuine rather than unread:updated 2026-08-064 in-linksSL2TEvery row but the first three is unknown, and the announcement is the reason ratherupdated 2026-08-134 in-linksTernary Bonsai 2 27BPricing is unknown rather than absent: PrismML publishes no hosted endpoint for this checkpoint, and the Toge…updated 2026-09-198 in-linksWeatherNext 3Context window and Pricing are unknown because neither applies in the formupdated 2026-09-052 in-linksWeatherNext CyclonesContext window and Pricing are unknown because neither applies in the form the rowupdated 2026-08-072 in-links
Concepts
/36Adversarial DistillationAdversarial distillation is the extraction of a model's capability from itsupdated 2026-10-017 in-linksAgent Runtime ContainmentAgent runtime containment is the problem of bounding what an autonomous agentupdated 2026-09-307 in-linksAgentic Reinforcement LearningA paradigm in which an LLM agent learns via RL while interacting with an environment. Instead of a single res…updated 2026-10-0495 in-linksAgents (LLM Agents)Systems that place an LLM at their core as the controller to perform multi-step planning + tool use + environ…updated 2026-10-05158 in-linksAI AlignmentThe problem of ensuring that AI systems reliably pursue goals that are beneficial to humans, and not just pro…updated 2026-10-0264 in-linksAI Control RoadmapA framework for securing AI systems at the system and infrastructure level — going beyond model-level alignme…updated 2026-08-0121 in-linksAI for MathematicsThe use of language models to produce new mathematical results — not to tutor, notupdated 2026-10-0418 in-linksAI GovernanceGovernance frameworks — legal, voluntary, and technical — that determine how frontier AI models are developed…updated 2026-10-0440 in-linksAI-Enabled CyberattacksThe use of AI models — either as intelligent assistants or fully autonomous agents — to conduct offensive cyb…updated 2026-10-0236 in-linksClaude Managed AgentsA cloud-hosted agent execution layer that separates agent logic (what Claude decides) from agent runtime (orc…updated 2026-09-1112 in-linksConceptual Reasoning Index (CRI)A composite benchmark that scores a model's conceptual reasoning — theupdated 2026-08-155 in-linksContent Provenance (AI output marking)Making a model's output identifiable as machine-generated after it has leftupdated 2026-10-0114 in-linksContext CompactionCompaction is what an agent does when a task outgrows its context window: itupdated 2026-09-189 in-linksEmbedded EvaluationEmbedded evaluation is third-party safety assessment performed from insideupdated 2026-09-2612 in-linksEmbodied AgentsAgents that act through perception and actuation in the physical world or in simulated environments. Unlike p…updated 2026-10-0232 in-linksEval Environment ContainmentEval environment containment is the problem of guaranteeing that a model beingupdated 2026-09-2835 in-linksEval Harness ConfigurationThe harness is the scaffolding around a model during a benchmark run: how context is carriedupdated 2026-10-05165 in-linksFrontier PacingThe proposition that the international system should build, in advance, the technical andupdated 2026-10-0236 in-linksGoogle ADK (Agent Development Kit)Google's Agent Development Kit (ADK) is an open-source, code-first toolkit for building, evaluating, and depl…updated 2026-07-017 in-linksGRAM — Gradient-Routed Auxiliary ModulesGRAM (Gradient-Routed Auxiliary Modules) is a modular pretraining architecture developed by Anthropic that is…updated 2026-07-212 in-linksLLM Knowledge Bases (LLM-curated personal wikis)A pattern for maintaining a continuously accumulating, structured personal or team knowledge base using an LL…updated 2026-08-318 in-linksMCP — Model Context ProtocolAn open protocol that standardizes how AI applications connect to external tools, data sources and services. …updated 2026-10-0416 in-linksMechanistic InterpretabilityMechanistic interpretability is the research program of reverse-engineering what specific internal computatio…updated 2026-10-0524 in-linksMHS — Model Hardware StandardA specification, announced by entities/anthropic on 2026-08-27 as aupdated 2026-08-295 in-linksMilitary and Intelligence Capability EvalsEvaluations that measure how well a model performs the specific technicalupdated 2026-09-143 in-linksModel RoutingChoosing which model answers which request — or which step of a request —updated 2026-09-2425 in-linksOpen-Weights Policy FightThe 2026 policy dispute over whether openly released model weights should be restricted, and on what grounds.…updated 2026-10-0466 in-linksPost-Training ScalingThe claim that the next increment of model capability is bought afterupdated 2026-10-0136 in-linksPreparedness Frameworkentities/openai's internal policy for deciding what a model is allowed to be —updated 2026-09-0622 in-linksR&D Automation IndexA measurement instrument that reports what fraction of a frontier lab's own AI research and development is pe…updated 2026-09-2914 in-linksReasoning ModelsA family of LLMs that explicitly model the reasoning process itself. They allocate test-time compute to reaso…updated 2026-08-1257 in-linksSafety CasesA safety case is a comprehensive, structured, evidence-based argument that aupdated 2026-09-309 in-linksSafety Monitoring and Data RetentionWhether a frontier lab must hold customer prompts and outputs in order toupdated 2026-09-0918 in-linksSoftware 3.0A taxonomy of software development paradigms presented by Andrej Karpathy at Sequoia Ascent 2026. A new era i…updated 2026-07-299 in-linksTest-Time Compute (Inference-Time Compute Scaling)Any technique that improves output quality by allocating additional compute at inference time. The trained mo…updated 2026-10-0555 in-linksWorld ModelsA world model is a learned model of an environment's dynamics: given a stateupdated 2026-09-308 in-links
People
/10Andrej KarpathyFormer Tesla AI director, founding member of OpenAI. Joined the entities/anthropic pretraining team on 2026-0…updated 2026-07-2710 in-linksChris OlahAnthropic co-founder. Pioneer of mechanistic interpretability — the research program of reverse-engineering w…updated 2026-07-309 in-linksJason WeiAI researcher and co-creator of chain-of-thought (CoT) prompting — one of the most influential techniques in …updated 2026-07-224 in-linksJeff DeanGoogle's Chief Scientist until August 2026, and one of the two engineers (with Sanjayupdated 2026-08-072 in-linksJim FanNVIDIA Senior Research Scientist. A leading researcher in Embodied AI / Foundation Agent. Leads Project GR00T…updated 2026-06-2710 in-linksJohn JumperComputational biologist; Nobel Laureate in Chemistry (2024); co-creator of AlphaFold at Google DeepMind. Anno…updated 2026-06-246 in-linksNoam ShazeerResearch scientist and engineer; VP Engineering at Google DeepMind; co-lead of the Gemini AI models; co-autho…updated 2026-06-245 in-linksPaul ChristianoAlignment researcher, now a governance figure at two institutions at once.updated 2026-09-101 in-linksSam AltmanCEO of entities/openai. Appears in this wiki less as a builder than as theupdated 2026-08-296 in-linksYann LeCunComputer scientist, AI pioneer, and one of the three "Godfathers of Deep Learning" (alongside Hinton and Beng…updated 2026-07-232 in-links
Papers
/2072028: Two Scenarios for Global AI Leadership — AnthropicAnthropic's policy essay argues that the US-China frontier AI gap will be decided by 2028, primarily through …updated 2026-05-183 in-linksA Case Study on Emergent Cheating and Whistleblowing in Autonomous Research SwarmsGoogle DeepMind ran 100 Gemini 3.1 Pro agents on 71 Lean 4 conjecturesupdated 2026-09-084 in-linksA Mechanistic View of Authority Hierarchy in LLM SycophancyHint a model toward a wrong answer and attribute the hint to a more seniorupdated 2026-10-022 in-linksAchieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified ScalingA recipe for converting an already post-trained reasoning backbone into anupdated 2026-08-144 in-linksAgensh: Scaling Organizational Intelligence to 1,024 AgentsA multi-agent harness with no central orchestrator — workers claim their ownupdated 2026-09-256 in-linksAgent Data Injection Attacks are Realistic Threats to AI AgentsA new attack class — Agent Data Injection (ADI) — exploits agents' trust in metadata (resource identifiers, t…updated 2026-07-143 in-linksAgent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528)Names and formalises harnessed agentic RL — the regime where the deploy-timeupdated 2026-08-206 in-linksAgent-Editing World Model: Rethinking World Modeling for LLM AgentsLanguage world models for agents predict environment observations; this paperupdated 2026-09-267 in-linksAgentGrad: Intervention-guided Prompt Optimization for Multi Agent SystemsCredit assignment in a multi-agent system, done by intervention rather than byupdated 2026-09-112 in-linksAgentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310)Argues that evolution strategies, not RL, are the right optimiser forupdated 2026-08-205 in-linksAgentic Misalignment in Summer 2026Follow-up to the 2025 blackmail experiment series. Catalogs four new agentic misalignment failure modes acros…updated 2026-07-213 in-linksAgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-CallingLLM judges are used to score agentic tool-calling systems, and nobody had checkedupdated 2026-09-033 in-linksAgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale (arXiv:2608.20634)Instead of building an environment for a task, AgentMercury instantiates aupdated 2026-08-266 in-linksAgora: Git as Shared Memory for Collective AutoResearchStores autonomous-research state as an append-only DAG of Git commits, so every claim is a commit anyone can …updated 2026-09-193 in-linksAI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (arXiv:2608.12307)A stronger model builds the harness; a weaker model runs inside it — and theupdated 2026-08-154 in-linksAn Empirical Study of Harness Design for Coding AgentsHolds a coding agent's execution loop fixed and varies three components — planning, action space, context man…updated 2026-09-193 in-linksAn Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsAn open-model pipeline that scored 30 out of 42 at IMO 2026 — the gold-medalupdated 2026-09-132 in-linksAn OpenAI model has disproved a central conjecture in discrete geometryAn OpenAI general-purpose reasoning model disproved the Erdős unit distance conjecture (planar unit distance …updated 2026-05-243 in-linksApodex 1.1: Scaling Agentic Intelligence for Complex Work (arXiv:2608.23283)entities/apodex's second system paper defines "working capability" —updated 2026-08-265 in-linksApodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (arXiv:2608.11341)Proposes evaluating AI on real-world problems that do not arrive in executableupdated 2026-08-185 in-linksAREX: Towards a Recursively Self-Improving Agent for Deep ResearchAn agent framework that recursively improves its own research pipelines using self-evaluated quality signals,…updated 2026-07-264 in-linksArgo-Bench: Evaluating Data Agents on Enterprise-Scale WorkflowsA 235-table, 7.5-billion-row simulated ERP warehouse where the agent is graded onupdated 2026-10-052 in-linksASI-Bench: At the Dawn of Artificial Superintelligence (arXiv:2608.17271)60 project-level research tasks across 11 scientific domains, built byupdated 2026-08-206 in-linksASPIRE: Agentic Skills Discovery for RoboticsA continual learning system (NVIDIA GEAR Lab + UMich/UIUC/Berkeley/CMU) that lets robots autonomously write a…updated 2026-07-072 in-linksAspire: Can Models Self-Evolve from Vague Goals?A benchmark that supplies only a natural-language capability goal — "become aupdated 2026-09-045 in-linksAtria Dawn: The Dawn of Agentic SuperintelligenceA model release with a labour study attached, and the study is the part worthupdated 2026-09-162 in-linksAutomated Weak-to-Strong Researcher (AAR)On an alignment research problem (weak-to-strong supervision), Anthropic's 9 AI agents achieved 97% PGR in th…updated 2026-05-319 in-linksAutonomous Mathematical Discovery in an Open-World Multi-Agent EnvironmentAgents from different model families were put in a shared open-worldupdated 2026-08-312 in-linksAutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces (arXiv:2608.23041)Harness improvement recast as an offline learning problem: diagnose failureupdated 2026-08-285 in-linksBDH-CQ: In-Context Learning with Recurrent Latent Reasoning (arXiv:2608.09888)A 150M-parameter model reasons by iterating in latent space instead ofupdated 2026-08-122 in-linksBeneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMsA chain of thought visibly does different kinds of work — formulating theupdated 2026-09-093 in-linksBeyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development (arXiv:2608.13417)Evaluates seven frontier models on 36 long-horizon AI-R&D tasks withupdated 2026-08-186 in-linksBeyond Solver Verdicts: Generative Reward Models for AutoformalizationA solver saying "proved" does not mean the thing it proved is the thing youupdated 2026-09-141 in-linksBeyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311)RL post-training controls drift with an action-side Policy-KL term, and thatupdated 2026-08-264 in-linksCapable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier ModelsUses a custom tool registered through a standard API feature to makeupdated 2026-09-252 in-linksChain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are DeliveredCoT monitoring assumes the reasoning trace records what actually shaped theupdated 2026-09-022 in-linksClawGym II: Exploring Black-Box RL on Agent Harness (arXiv:2608.16798)Trains a model through a harness it does not control, by placing a servingupdated 2026-08-197 in-linksCo-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL (arXiv:2608.17253)Trains several parameter-independent models simultaneously with RL, eachupdated 2026-08-213 in-linksCOBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill OptimizationAgent skills are cheap to write and expensive to evaluate, so this paper spendsupdated 2026-09-151 in-linksCodeMidas: Scaling Agentic Coding RL Environments from Code ItselfAn agentic pipeline that turns existing source code — and nothing else — intoupdated 2026-09-223 in-linksCoding Agents for Generalized Task and Motion Planning ProblemsThree production coding agents — Claude Code (Opus 5) and Codex onupdated 2026-09-284 in-linksConfidence Comes from Experience: Experiential Confidence Estimation from Reasoning to AgentsEvery existing confidence estimator reads only the current inference —updated 2026-09-212 in-linksContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RLLong-horizon agents accumulate an ever-growing working context. Prior work lets aupdated 2026-09-018 in-linksContinual Learning Mechanisms Compose for Long-Horizon MemorizationNo single continual-learning mechanism holds up over 100 sequentialupdated 2026-09-181 in-linksCook and Clean Together: Teaching Embodied Agents for Parallel Task Execution (GRANT)Embodied agents execute instructions serially even when the physical world permits overlap —updated 2026-07-311 in-linksDarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545)The model is frozen and the harness evolves. DarwinX runs a population ofupdated 2026-08-1611 in-linksDataPrep-Bench: Benchmarking LLMs as Training Data PreparatorsThe first benchmark that scores an LLM on preparing training data end to end — both buildingupdated 2026-07-30Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning (arXiv:2608.18746)JEPA-style latent world models plan by taking Euclidean distance to a goal latentupdated 2026-08-237 in-linksDecoding Looped Transformers Better for (Almost) FreeA looped Transformer already computes a decodable prediction at every recurrentupdated 2026-10-053 in-linksDemystifying Agent Skills: Why They Work — Until They Don't (arXiv:2608.14036)Isolates what a skill actually does for an LLM agent and finds it is not whatupdated 2026-08-207 in-linksDFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data (arXiv:2608.13517)A 1B-parameter model on the Hierarchical Reasoning Model (HRM)updated 2026-08-181 in-linksDiffuse AI Control on Fuzzy TasksA red-teaming framework for training interventions against diffuse threats on fuzzy tasks —updated 2026-07-313 in-linksDoes On-Policy Distillation Really Distill? From Noisy Teacher to Self-ImprovementOn-policy distillation is supposed to work by giving a student dense token-levelupdated 2026-09-023 in-linksEliciting Weak-to-Strong Generalization with On-Policy Reverse DistillationConventional distillation makes the teacher the target, which hands the studentupdated 2026-09-103 in-linksEngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?1,301 expert-curated tasks across 6 engineering domains and 26 professionalupdated 2026-10-011 in-linksENPIRE: Agentic Robot Policy Self-Improvement in the Real WorldA fleet of 8 real robots autonomously runs its own research loop — reading papers, proposing hypotheses, runn…updated 2026-06-276 in-linksEnvHarness: Awakening Static Worlds for Agent Learning (arXiv:2608.19880)A programmable plug-in layer that wraps a static training environment andupdated 2026-08-227 in-linksEvery Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647)A controlled study that varies one generalization factor at a time findsupdated 2026-08-254 in-linksEvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent HarnessesAgents that rewrite their own prompts, tools and harnesses can improve — and canupdated 2026-09-024 in-linksExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien WorldsEvaluating scientific exploration has two hard problems — verifying a genuinelyupdated 2026-09-275 in-linksFACET: Preserving Source Intent and Executable State in Terminal Task Synthesis (arXiv:2608.18580)A framework for synthesizing terminal-agent training tasks where the four coupledupdated 2026-08-224 in-linksFalse Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search AgentsA self-evolving search agent that generates its own questions and answers themupdated 2026-10-042 in-linksFeyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber ModelsSeven people trained open-weight models to a leading agentic cyber capability,updated 2026-09-152 in-linksFlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving (arXiv:2608.19758)Turns the authors' earlier FlashPrefill prototype into something a serving stack canupdated 2026-08-231 in-linksFlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning ExperienceA reasoning model can improve from its own rollouts, but the loop is fragile fromupdated 2026-09-093 in-linksFlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills (arXiv:2607.21596)An agent that builds a workflow at inference time normally throws it away whenupdated 2026-08-256 in-linksFM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (arXiv:2608.18423)An LLM agent runs a football club for 20 in-game years — ~340–400 decisionupdated 2026-08-225 in-linksFreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157)An edge-native MoE serving system that treats a personal machine as a unifiedupdated 2026-08-202 in-linksFrontierChallenge: Evaluating Scientific Workflow Completion (arXiv:2608.24979)A cross-domain benchmark of 300 end-to-end scientific workflows, of which 97updated 2026-08-285 in-linksFull-bandwidth transformer (arXiv:2608.08888)Widens the vertical channel between decoding steps: instead of only theupdated 2026-08-174 in-linksGeneralized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-ImprovementProposes Generalized Agent Iteration (GAI), a formal framework that makesupdated 2026-09-183 in-linksGenerative design of bacteriophages with genome language models (Science, DOI 10.1126/science.aec2657)A genome language model (Evo 2) wrote complete bacteriophage genomes fromupdated 2026-08-083 in-linksGroupwise Agentic Grading and Advantage Redistribution for Code Agent RLBinary test-pass rewards make GRPO blind to code quality: every trajectoryupdated 2026-09-304 in-linksHappyWorld-BenchThe first instrument this wiki holds that evaluates world models on whether theupdated 2026-09-272 in-linksHarmProfile: Characterizing Harmful Distributions in Frontier LLMs (arXiv:2608.14577)Treats harmful generation as an object of analysis rather than an attackupdated 2026-08-202 in-linksHarness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents (arXiv:2608.15008)Runs every common memory substrate for LLM agents through one unified harnessupdated 2026-08-205 in-linksHarness-of-Harness: Multi-Day Autonomous Software Development with Continual ImprovementHoH is a framework that wraps existing coding-agent harnesses and organisesupdated 2026-09-034 in-linksHarness-Zero: Harness Distillation via Agent-as-HarnessTakes the gains a specialized agent harness produces and trains them into theupdated 2026-09-238 in-linksHarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?A benchmark that shifts the unit of evaluation from task outputs to runnableupdated 2026-09-046 in-linksHierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses (arXiv:2608.08466)The harness — the executable scaffold around the model — is normally a fixedupdated 2026-08-259 in-linksHow Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review (arXiv:2608.08975)Holds the reported scientific content fixed and varies only how it isupdated 2026-08-176 in-linksHow Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905)Runs 8 harness-model combinations over 100 research tasks, producingupdated 2026-08-1911 in-linksHunyuan-A13B Technical Reportentities/tencent's technical report for Hunyuan-A13B, an open-source MoEupdated 2026-09-272 in-linksImprint Reader: From Weight-Update Readout to Behavioral InterventionTraining leaves parameter-level traces, and current models cannot say what thoseupdated 2026-09-303 in-linksImproving the matrix multiplication exponent with modern optimization and AlphaEvolve (arXiv:2608.16884)Improves the best known upper bound on the matrix multiplication exponent ωupdated 2026-08-191 in-linksIntern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (arXiv:2608.14290)Mobius-v0 splits a transformer into a globally shared Memory (FFN) holdingupdated 2026-08-182 in-linksIntern-S2-Preview: Scientific Agentic Foundation Model (arXiv:2608.13505)A 397B scientific agentic foundation model, trained through multimodalupdated 2026-08-173 in-linksIris: Climbing to the Search FrontierTwo open-source search agents — Iris-mini (35B-A3B) and Iris-proupdated 2026-09-083 in-linksJ-Zero: Unified Challenger–Solver–Judge Co-Evolution from Zero DataSelf-evolving models have made progress where an automatic verifier can score theupdated 2026-09-016 in-linksJEV-as-a-Judge: Accept When Confident, Escalate When UnsureThe first independent evaluation of models/jev this wiki holds. Aupdated 2026-09-244 in-linksJev-Mem: System-One-Controlled Agentic Memory for Efficient AI AgentsTakes the expensive autoregressive LLM off the critical path of memoryupdated 2026-09-243 in-linksJIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution (arXiv:2608.25593)A model trained to write agent harnesses. JIT-Agent formalises the harness asupdated 2026-08-285 in-linksJust-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM AgentsAgent memory systems decide what to keep at write time, before the futureupdated 2026-09-266 in-linksKnowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211)Names futile reasoning — expensive, semantically void reasoning produced onupdated 2026-08-183 in-linksLanguage Models Are "Insecure" ReportersHanded experiment logs containing a planted negative result that undermines theupdated 2026-10-013 in-linksLast Translation BenchmarkA benchmark built only from examples that break leading machine-translationupdated 2026-09-064 in-linksLearning to Discover Interesting MathematicsEvery AI-for-mathematics result this wiki holds is scored against problems humansupdated 2026-09-283 in-linksLearning to Solve Hard Problems in RL for LLMs by Never Giving UpRL post-training makes models better at what they were already good at, and theupdated 2026-09-17LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393)Trains a coding agent inside three unmodified production harnesses — OpenHandsupdated 2026-08-2112 in-linksLet's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts (arXiv:2608.20061)Sweeping for the optimal learning rate is prohibitive at trillion-token, MoEupdated 2026-08-251 in-linksLLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers (arXiv:2608.06867)The first attempt to make model routers comparable to each other: oneupdated 2026-08-161 in-linksLLMs are General Asynchronous AgentsThe read–think–reply loop is the assumption, and it is dropped. Modern agentsupdated 2026-10-011 in-linksLong-Horizon-Terminal-Bench (LHTB)A 46-task containerized terminal benchmark with dense reward grading; current best model achieves only 15.2% …updated 2026-07-151 in-linksLoopArena: Benchmarking Models as Runtime Controllers for Loop EngineeringWhen a coding agent fails a long task, nothing tells you whether the loopupdated 2026-09-014 in-linksLooped Language Models Improve Compositional Tool Calling (arXiv:2608.18171)Tests looped (recurrent-depth) language models on compositional tool use —updated 2026-08-211 in-linksMechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (arXiv:2608.12036)An agentic system that does mechanistic interpretability research on its own —updated 2026-08-163 in-linksMemory as Plans: World-Action Modeling with Memory-Grounded PlanningRobot policies are mostly Markovian and many real manipulation tasks are not.updated 2026-09-141 in-linksMemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use (arXiv:2608.20202)A benchmark for memory-induced cognitive traps: even faithfully recorded,updated 2026-08-228 in-linksMeta^n: Recursive Self-Improvement through Emergent Depth (arXiv:2608.24735)Self-improving systems cap out at about two levels of meta-depth, because aupdated 2026-08-284 in-linksMid-Harness: Scaling Actions Between Model and Harness for Terminal AgentsSpend test-time compute at the boundary between the model and the harness: sampleupdated 2026-10-044 in-linksModel Spec Midtraining: Improving How Alignment Training GeneralizesTrain the model on documents about its own spec before you train it onupdated 2026-10-012 in-linksMolt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement LearningNVIDIA NeMo's Molt is a compact, researcher-friendly PyTorch-native RL training framework for agentic scenari…updated 2026-07-282 in-linksMore than two thirds of the zeros of the Riemann zeta function lie on the critical lineAn unreleased research version of Claude raised the proven lower bound on theupdated 2026-08-113 in-linksNegative Self-Distillation: Learning to Reason by Avoiding FlawsInverts On-Policy Self-Distillation: instead of imitating a privileged teacher,updated 2026-09-132 in-linksNeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing HarnessRecursive self-improvement is claimed here as a training-data pipeline, not asupdated 2026-09-102 in-linksOmniScientist: An Omni-Modal Omni-Discipline AI Scientist (arXiv:2608.13558)An AI scientist that works from heterogeneous raw evidence — images, signals,updated 2026-08-211 in-linksOn the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training StabilityAlibaba's own architecture report for Qwen3.8-Flash-Next: 125B sparse MoE,updated 2026-09-022 in-linksOn-Policy or Off-Policy Learning? A Systematic Study of Distillation DynamicsA controlled study that varies rollout policy, token-level KL direction andupdated 2026-10-041 in-linksOne Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (arXiv:2608.19741)A sandbox with isolated MCP-compatible tool sessions, complete executionupdated 2026-08-269 in-linksOne Symptom, Three Levers: A Critical Review of On-Policy Self-DistillationOn-policy distillation trains a model on its own generations while a largerupdated 2026-09-096 in-linksOpenAI Parameter Golf — What It Taught UsOpenAI ran a community ML challenge (16 MB model, 10 min training, 8×H100s). Key finding: AI coding agents ha…updated 2026-05-183 in-linksOpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution (arXiv:2608.00677)Red-team the environment, not the prompt. OpenART is an arena of 10,000+updated 2026-08-142 in-linksPACT: From Credit Assignment to Critic AlignmentToken-level credit in LLM RL "lacks a generally accepted mathematicalupdated 2026-09-262 in-linksParaTempo: Efficient Parallel Reasoning via Temporal Confidence (arXiv:2608.16425)Parallel reasoning buys accuracy by exploring several solution paths, and pays forupdated 2026-08-251 in-linksPARSER: Read in Parallel, Reason in Depth for Long-Context LLM AgentsSequential memory agents couple how far they have read to how deeply they haveupdated 2026-09-121 in-linksPILOT in the Loop: Live Self-Improvement for Long-Horizon Agents (arXiv:2608.26530)A supervisor-worker harness that updates itself during a run rather thanupdated 2026-08-295 in-linksPositive Alignment: Artificial Intelligence for Human FlourishingA 16-author collaborative paper from Oxford, DeepMind, Anthropic, and others. It argues that today's alignmen…updated 2026-05-313 in-linksPost-Training Leaves Behavioral Shadows on Unrelated DecisionsA capability can be transferred between models using one word of teacher outputupdated 2026-09-305 in-linksPrime Agent: A Self-Improving RLM Harness (arXiv:2608.23552)An open-source harness for long-horizon evaluation and coding-agent work,updated 2026-08-266 in-linksProcedural Graphs: Self-Evolving Execution Structures for LLM AgentsA knowledge graph stores (entity, relation, entity) for what-is questions; aupdated 2026-09-102 in-linksProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE TasksA coding-agent benchmark where the spec is a working application rather than anupdated 2026-09-215 in-linksProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM AgentsThe failure it names is a claim that is true somewhere in the evidence andupdated 2026-10-012 in-linksQuantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs (arXiv:2608.20953)Serving cheaply increasingly means shipping a model that is both structurallyupdated 2026-08-261 in-linksQwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI AgentsA foundation GUI agent from Tongyi MAI (entities/alibaba) that unifies mobile,updated 2026-08-032 in-linksR³-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets (arXiv:2608.16033)Gives a model six problems and one budget between them, and measures how wellupdated 2026-08-191 in-linksRealSWE: A Compositional Evaluation of Coding Agents under Realistic User RequestsMeasures the gap between how SWE-bench problems are written and how realupdated 2026-09-053 in-linksRecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use AgentsComputer-use and coding have been measured separately; this benchmark makes anupdated 2026-09-222 in-linksRecursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses (arXiv:2608.24876)Recuris splits agent memory in two — Working Memory tracking task progress,updated 2026-08-284 in-linksRelic: From Multi-Agent Collaboration to Persistent Organizational CapabilityMulti-agent systems resolve a conflict in conversation and then lose theupdated 2026-09-305 in-linksRepo-To-Skill: Distilling GitHub Repositories Into AI4AI SkillsNames the layer a research agent is missing — operational knowledge, "theupdated 2026-09-044 in-linksRepo0: Design-Driven Zero-to-All Code Generation (arXiv:2608.19854)Most coding agents assume a repository architecture already exists. Repo0 targetsupdated 2026-08-233 in-linksRethinking Critic Learning in PPO: Understanding and Mitigating Value FlatteningNames a systematic failure mode in PPO critics for LLM training — Valueupdated 2026-09-202 in-linksRethinking On-Policy Distillation of Large Language Models II: One Training ExampleTrains on-policy distillation on a single query and recovers most ofupdated 2026-09-051 in-linksRewardVerse: Rubric-Guided Policy Optimization for Video Reward ModelingVideo reward models are asked to compress "is this video good" into one scalar, andupdated 2026-09-282 in-linksRing-Zero: Scaling Zero RL to a Trillion Parameters for Emergent ReasoningFirst demonstration of RLVR (RL with Verifiable Rewards) scaling to 1 trillion parameters — achieves 84.2% on…updated 2026-07-196 in-linksRound-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors (arXiv:2608.00675)Train one latent diffusion model that can step a dynamical system forwards or backwardsupdated 2026-08-07RRSI: Regularized Recursive Self-Improvement of Agent HarnessesAgent harnesses that improve themselves overfit the tasks they are evolvedupdated 2026-09-232 in-linksRufus-Air: An Open LLM Post-Training RecipeAn open and reproducible post-training recipe on GLM-4.5-Air-Baseupdated 2026-09-275 in-linksSAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?A benchmark that asks whether an agent can do the auditing work, not just theupdated 2026-09-111 in-linksScaling Automatic Research Agents via World ModelsThe bottleneck in training a research agent is not the model, it is the sandbox.updated 2026-09-121 in-linksScaling Properties of Same-Family On-Policy DistillationPower laws for on-policy distillation, and the headline result is that a smallupdated 2026-10-012 in-linksScaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B AgentA 35B MoE model (Agents-A1) matches 1-trillion-parameter models on agentic benchmarks by scaling the agent ho…updated 2026-07-015 in-linksSchrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?Proposes SchrodingerRepo, which rebuilds the test repository at evaluationupdated 2026-09-255 in-linksScores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research AgentsA protocol that treats a research agent's claimed discovery as a hypothesis to beupdated 2026-09-114 in-linksSecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation (arXiv:2608.21500)Defensive fine-tuning against prompt injection has been failing because its trainingupdated 2026-08-286 in-linksSEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement LearningSEED converts an LLM agent's own completed trajectories into natural-language "hindsight skills," then distil…updated 2026-07-195 in-linksSelect, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMsA long-video model cannot look at every frame — an hour sampled once per secondupdated 2026-09-082 in-linksSelf-Distilled Agentic Reinforcement LearningSDAR adds token-level distillation guidance to reinforcement learning forupdated 2026-08-146 in-linksSemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation (arXiv:2608.18565)An agent harness for industrial PLC code whose defining rule is that a task isupdated 2026-08-218 in-linksSemComp-Bench: Benchmarking Semantic Task Completion in Video Generation (arXiv:2608.17426)Reformulates video generation as an outcome-oriented task: success requires theupdated 2026-08-231 in-linksShieldstral (arXiv:2607.25857)A 3B multimodal safety classifier that takes its moderation policy as naturalupdated 2026-08-062 in-linksSimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning (arXiv:2608.14277)On-policy distillation across tokenizers: align only the tokens occupyingupdated 2026-08-183 in-linksSkill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving SkillsSkill-SP makes agent skills the unit that a self-play loop evolves, using each skill's narrow, verifiable exe…updated 2026-07-297 in-linksSkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback (arXiv:2608.13120)Argues the binding constraint on evolving agent skills is not editing capabilityupdated 2026-08-224 in-linksSkillForge: Self-Distilling Agents for Project-Specific Issue Resolution (arXiv:2608.18933)A self-distillation framework that acquires project-specific knowledge from theupdated 2026-08-222 in-linksSLEIGHT-Bench: Finding Blind Spots in AI MonitorsA benchmark of 40 synthetic transcripts across 11 categories, each one a coding agent doingupdated 2026-07-313 in-linksSMELT: Scaling Laws for Compute-Matched MoE Looped TransformersLooped Transformers gain effective depth by iterating a shared block, but theupdated 2026-09-031 in-linksSoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation (arXiv:2608.18701)A visuo-tactile dataset and closed-loop benchmark for deformable-objectupdated 2026-08-235 in-linksSolipsistic Superintelligence is Unlikely to be CooperativeCurrent RL/RLHF training treats the world as a stationary, exogenous environment. Deployed systems break this…updated 2026-06-066 in-linksSPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv:2608.19197)One LLM plays two roles: an Environment Designer that writes completeupdated 2026-08-218 in-linksSpatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence (arXiv:2608.12743)Asks whether a frozen VLM can improve its spatial reasoning with noupdated 2026-08-173 in-linksSpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party DialogueGeneral-purpose agent memory retrieves content and loses who said it aboutupdated 2026-09-272 in-linksStateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (arXiv:2608.15089)An agent runtime that changes nothing about the model weights and reportsupdated 2026-08-1910 in-linksStealing Reasoning Traces from Proprietary LLM APIs (arXiv:2608.09867)Providers hide chain-of-thought by returning it to the client as an encryptedupdated 2026-08-124 in-linksSteering Geometry: Validating Human Value Geometry in LLM Steering SpaceTwo families of steering method work equally well and only one of them isupdated 2026-09-102 in-linksSWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? (arXiv:2608.23564)20 whole-repository migrations, evaluated in three stages so that "the testsupdated 2026-08-292 in-linksSWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering AgentsThe benchmark 24 model pages on this wiki quote has been audited by its ownupdated 2026-09-116 in-linksSWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? (arXiv:2608.19799)A repository-level benchmark of 119 tasks from 98 GitHub repositories across 20updated 2026-08-228 in-linksT1: Terminal Agent Reinforcement Learning for Long-Horizon TasksA 122B Mixture-of-Experts model trained by RL to drive a real shell in a cloudupdated 2026-09-122 in-linksTerminal-Universe: Turning Agent Trajectories into Scalable Terminal EnvironmentsReconstructs executable environments out of the agent trajectories that alreadyupdated 2026-09-055 in-linksThe Embedder's Dilemma: LLMs Are Better, but at What Cost? (arXiv:2608.12875)Should an LLM replace a text-embedding pipeline? On aggregate the two paradigmsupdated 2026-08-251 in-linksThe Handoff Tax: Continuing Non-Native Trajectories in LLM Agents (arXiv:2608.24358)Measures what it costs when one model picks up a long agent trajectory anotherupdated 2026-08-291 in-linksThe Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement LearningLLM RL implementations use separate inference and training engines for efficiency, creating a systematic trai…updated 2026-07-175 in-linksThe More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning (arXiv:2608.14229)Popular facts are memorised more deeply in pretraining and resist removal longer,updated 2026-08-231 in-linksThe Tasteful Agent: Measuring and Improving Taste in Long-Horizon TasksBuilds Taste-Bench, which measures not whether an agent finishes aupdated 2026-09-242 in-linksThinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See (arXiv:2608.17744)Three frontier MoE models fine-tuned to reason in Greek show almost no accuracyupdated 2026-08-233 in-linksThought-Level Beam Search for Reasoning (arXiv:2608.08020)Reframes test-time compute as a allocation problem rather than a budgetupdated 2026-08-175 in-linksTraining Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report (arXiv:2608.15763)The harness cluster's mirror image. Whereupdated 2026-08-293 in-linksTraining Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929)A data-free, white-box test for whether two model checkpoints share ancestry,updated 2026-08-213 in-linksTransformers Stop Thinking Too Early, and a Tiny LoRA Fixes ItThirteen base models can follow only 1.4–3.6 links of a reference chain inupdated 2026-10-053 in-linksTTPO: Test-Time Policy Optimization (arXiv:2608.27448)Post-training without labels, by treating agreement and disagreement with aupdated 2026-08-293 in-linksUnderstanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO (arXiv:2608.27351)Evolution Strategies — perturb the whole parameter vector, score theupdated 2026-08-293 in-linksUnlocking Lossless Speedups in LLMs via Discrete DiffusionUno is an autoregressive model that draws several tokens at once. Theupdated 2026-09-091 in-linksUsing Grounded Theory for Agent Behavior Analysis at ScaleAutoTraceGT automates grounded theory — a six-decade-old qualitativeupdated 2026-09-084 in-linksVentor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs (arXiv:2608.16391)A black-box audit of what a hosted API actually serves you, requiring noupdated 2026-08-193 in-linksVerifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved MechanismsEnvironment-generation pipelines build the world first and bolt the reward onupdated 2026-09-274 in-linksWeak-to-Strong Generalization via Direct On-Policy DistillationRun RL on a cheap small model; transfer only the RL-induced policy delta (not the full policy) to a large mod…updated 2026-07-167 in-linksWHALE: A Simple Recipe for Joint Harness-Weight OptimizationAlternates updating the weights under the current harness with searchingupdated 2026-09-056 in-linksWhat LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two FleetsSix months of autonomous trading agents running with real money, measured as aupdated 2026-09-101 in-linksWhen Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token AnalysisAgents convert compute into quality faster than brute-force sampling at first,updated 2026-09-164 in-linksWhen EOS Tokens Disagree: Understanding Length Inflation in On-Policy DistillationStudent models distilled on-policy from post-trained teachers produceupdated 2026-09-202 in-linksWikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill EvolutionAgents that discover their own skills from experience keep the reasoning thatupdated 2026-08-317 in-linksWould this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual ExperimentsAnthropic proposes counterfactual simulatability as the test of an explanationupdated 2026-08-272 in-linksYour Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMsLinearly combine the inputs from two distinct text streams and the modelupdated 2026-09-261 in-linksZetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence (arXiv:2608.16590)Applies harness self-evolution to physical execution: Zetta evolves code-basedupdated 2026-08-217 in-linksZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic SearchA 7B dense model argues that a small model should stop trying to remember theupdated 2026-09-171 in-linksτ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation (arXiv:2608.16885)Hierarchical vision-language-action models pick the next subtask in a singleupdated 2026-08-242 in-links