$ tree wiki/
위키
한국어로 번역된 433개 페이지입니다. 번역이 원문과 일치하는 문서만 여기 실립니다 — 나머지는 영어 위키에서 보실 수 있습니다.
동향 · 24
- 주간 종합 — W39 (2026-09-21 → 2026-09-27)
- 2026-W40
- 2026년 8월 — 월간 다이제스트
- 2026년 7월 — 월간 다이제스트
- 월간 종합 — 2026년 6월
- 2026년 9월 — 월간 다이제스트
- 주간 종합 — 2026-W21 (2026-05-11 ~ 2026-05-17)
- 주간 종합 — 2026-W22 (2026-05-18 ~ 2026-05-24)
- 주간 종합 — 2026-W23 (2026-05-25 ~ 2026-05-31)
- 주간 종합 — 2026-W24 (2026-06-01 ~ 2026-06-07)
- 주간 종합 — 2026-W25 (2026-06-08 ~ 2026-06-21)
- 주간 종합 — 2026-W26 (2026-06-22 ~ 2026-06-28)
- 주간 종합 — 2026-W27 (2026-06-29 ~ 2026-07-05)
- 주간 종합 — 2026-W28 (2026-07-06 ~ 2026-07-12)
- 주간 종합 — 2026-W29 (2026-07-13 ~ 2026-07-19)
- 주간 종합 — W30 (2026년 7월 20~26일)
- 주간 종합 — W31 (2026년 7월 27일 – 8월 2일)
- 주간 종합 — W32 (2026-08-03 → 2026-08-09)
- 주간 종합 — W33 (2026-08-10 → 2026-08-16)
- 주간 종합 — W34 (2026-08-17 → 2026-08-23)
- 주간 종합 — W35 (2026-08-24 → 2026-08-30)
- 주간 종합 — W36 (2026-08-31 → 2026-09-06)
- 주간 종합 — W37 (2026-09-07 → 2026-09-13)
- 주간 종합 — W38 (2026-09-14 → 2026-09-20)
논문 · 207
- 2028: Two Scenarios for Global AI Leadership — Anthropic
- A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
- A Mechanistic View of Authority Hierarchy in LLM Sycophancy
- Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
- Agensh: Scaling Organizational Intelligence to 1,024 Agents
- 에이전트 데이터 주입 공격은 AI 에이전트에 대한 현실적 위협이다
- Agent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528)
- Agent-Editing World Model: Rethinking World Modeling for LLM Agents
- AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
- Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310)
- Agentic Misalignment in Summer 2026
- AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
- AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale (arXiv:2608.20634)
- Agora: Git as Shared Memory for Collective AutoResearch
- AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (arXiv:2608.12307)
- An Empirical Study of Harness Design for Coding Agents
- An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
- An OpenAI model has disproved a central conjecture in discrete geometry
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work (arXiv:2608.23283)
- Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (arXiv:2608.11341)
- AREX: Towards a Recursively Self-Improving Agent for Deep Research
- Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
- ASI-Bench: At the Dawn of Artificial Superintelligence (arXiv:2608.17271)
- ASPIRE: 로봇을 위한 에이전트 스킬 발견
- Aspire: Can Models Self-Evolve from Vague Goals?
- Atria Dawn: The Dawn of Agentic Superintelligence
- Automated Weak-to-Strong Researcher (AAR)
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
- AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces (arXiv:2608.23041)
- BDH-CQ: In-Context Learning with Recurrent Latent Reasoning (arXiv:2608.09888)
- Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development (arXiv:2608.13417)
- Beyond Solver Verdicts: Generative Reward Models for Autoformalization
- Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311)
- Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
- Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
- ClawGym II: Exploring Black-Box RL on Agent Harness (arXiv:2608.16798)
- Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL (arXiv:2608.17253)
- COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
- CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
- Coding Agents for Generalized Task and Motion Planning Problems
- Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
- ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
- Continual Learning Mechanisms Compose for Long-Horizon Memorization
- Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution (GRANT)
- DarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545)
- DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
- Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning (arXiv:2608.18746)
- Decoding Looped Transformers Better for (Almost) Free
- Demystifying Agent Skills: Why They Work — Until They Don't (arXiv:2608.14036)
- DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data (arXiv:2608.13517)
- Diffuse AI Control on Fuzzy Tasks
- Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
- Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
- EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?
- ENPIRE: 실세계 에이전트 로봇 정책 자기개선
- EnvHarness: Awakening Static Worlds for Agent Learning (arXiv:2608.19880)
- Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647)
- EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
- ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
- FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis (arXiv:2608.18580)
- False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
- Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving (arXiv:2608.19758)
- FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
- FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills (arXiv:2607.21596)
- FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (arXiv:2608.18423)
- FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157)
- FrontierChallenge: Evaluating Scientific Workflow Completion (arXiv:2608.24979)
- Full-bandwidth transformer (arXiv:2608.08888)
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
- Generative design of bacteriophages with genome language models (Science, DOI 10.1126/science.aec2657)
- Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
- HappyWorld-Bench
- HarmProfile: Characterizing Harmful Distributions in Frontier LLMs (arXiv:2608.14577)
- Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents (arXiv:2608.15008)
- Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
- Harness-Zero: Harness Distillation via Agent-as-Harness
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses (arXiv:2608.08466)
- How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review (arXiv:2608.08975)
- How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905)
- Hunyuan-A13B Technical Report
- Imprint Reader: From Weight-Update Readout to Behavioral Intervention
- Improving the matrix multiplication exponent with modern optimization and AlphaEvolve (arXiv:2608.16884)
- Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (arXiv:2608.14290)
- Intern-S2-Preview: Scientific Agentic Foundation Model (arXiv:2608.13505)
- Iris: Climbing to the Search Frontier
- J-Zero: Unified Challenger–Solver–Judge Co-Evolution from Zero Data
- JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
- Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
- JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution (arXiv:2608.25593)
- Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
- Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211)
- Language Models Are "Insecure" Reporters
- Last Translation Benchmark
- Learning to Discover Interesting Mathematics
- Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
- LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393)
- Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts (arXiv:2608.20061)
- LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers (arXiv:2608.06867)
- LLMs are General Asynchronous Agents
- Long-Horizon-Terminal-Bench (LHTB)
- LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
- Looped Language Models Improve Compositional Tool Calling (arXiv:2608.18171)
- Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (arXiv:2608.12036)
- Memory as Plans: World-Action Modeling with Memory-Grounded Planning
- MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use (arXiv:2608.20202)
- Meta^n: Recursive Self-Improvement through Emergent Depth (arXiv:2608.24735)
- Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
- Model Spec Midtraining: Improving How Alignment Training Generalizes
- Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
- More than two thirds of the zeros of the Riemann zeta function lie on the critical line
- Negative Self-Distillation: Learning to Reason by Avoiding Flaws
- NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
- OmniScientist: An Omni-Modal Omni-Discipline AI Scientist (arXiv:2608.13558)
- On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
- On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
- One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (arXiv:2608.19741)
- One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
- OpenAI Parameter Golf — What It Taught Us
- OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution (arXiv:2608.00677)
- PACT: From Credit Assignment to Critic Alignment
- ParaTempo: Efficient Parallel Reasoning via Temporal Confidence (arXiv:2608.16425)
- PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
- PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents (arXiv:2608.26530)
- Positive Alignment: Artificial Intelligence for Human Flourishing
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions
- Prime Agent: A Self-Improving RLM Harness (arXiv:2608.23552)
- Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
- ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
- ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
- Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs (arXiv:2608.20953)
- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
- R³-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets (arXiv:2608.16033)
- RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
- RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
- Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses (arXiv:2608.24876)
- Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
- Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
- Repo0: Design-Driven Zero-to-All Code Generation (arXiv:2608.19854)
- Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
- Rethinking On-Policy Distillation of Large Language Models II: One Training Example
- RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling
- Ring-Zero: 창발적 추론을 위한 Zero RL의 1조 매개변수 규모 확장
- Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors (arXiv:2608.00675)
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Rufus-Air: An Open LLM Post-Training Recipe
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
- Scaling Automatic Research Agents via World Models
- Scaling Properties of Same-Family On-Policy Distillation
- 매개변수가 아닌 수평선을 확장: 35B 에이전트로 1조 매개변수 성능 달성
- Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?
- Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
- SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation (arXiv:2608.21500)
- SEED: 에이전트 강화학습을 위한 자기 진화 온폴리시 증류
- Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
- Self-Distilled Agentic Reinforcement Learning
- SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation (arXiv:2608.18565)
- SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation (arXiv:2608.17426)
- Shieldstral (arXiv:2607.25857)
- SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning (arXiv:2608.14277)
- Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
- SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback (arXiv:2608.13120)
- SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution (arXiv:2608.18933)
- SLEIGHT-Bench: Finding Blind Spots in AI Monitors
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
- SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation (arXiv:2608.18701)
- Solipsistic Superintelligence is Unlikely to be Cooperative
- SPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv:2608.19197)
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence (arXiv:2608.12743)
- SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue
- StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (arXiv:2608.15089)
- Stealing Reasoning Traces from Proprietary LLM APIs (arXiv:2608.09867)
- Steering Geometry: Validating Human Value Geometry in LLM Steering Space
- SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? (arXiv:2608.23564)
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? (arXiv:2608.19799)
- T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
- Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
- The Embedder's Dilemma: LLMs Are Better, but at What Cost? (arXiv:2608.12875)
- The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents (arXiv:2608.24358)
- The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
- The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning (arXiv:2608.14229)
- The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
- Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See (arXiv:2608.17744)
- Thought-Level Beam Search for Reasoning (arXiv:2608.08020)
- Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report (arXiv:2608.15763)
- Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929)
- Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
- TTPO: Test-Time Policy Optimization (arXiv:2608.27448)
- Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO (arXiv:2608.27351)
- Unlocking Lossless Speedups in LLMs via Discrete Diffusion
- Using Grounded Theory for Agent Behavior Analysis at Scale
- Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs (arXiv:2608.16391)
- Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
- 직접 온폴리시 증류를 통한 약한 모델에서 강한 모델로의 일반화
- WHALE: A Simple Recipe for Joint Harness-Weight Optimization
- What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
- When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
- When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
- WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
- Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
- Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
- Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence (arXiv:2608.16590)
- ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
- τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation (arXiv:2608.16885)
개념 · 36
- 적대적 증류 (Adversarial Distillation)
- 에이전트 런타임 격리
- 에이전틱 강화학습
- 에이전트 (LLM Agents)
- AI 정렬 (Alignment)
- AI 통제 로드맵
- 수학을 위한 AI (AI for Mathematics)
- AI 거버넌스
- AI 기반 사이버 공격
- Claude Managed Agents
- Conceptual Reasoning Index (CRI)
- 콘텐츠 프로버넌스 (AI 출력 마킹)
- 컨텍스트 압축 (Context Compaction)
- 상주 평가 (Embedded Evaluation)
- 체화 에이전트 (Embodied Agents)
- 평가 환경 격리
- 평가 하네스 설정
- 프런티어 속도 조절 (Frontier Pacing)
- Google ADK (Agent Development Kit)
- GRAM — Gradient-Routed Auxiliary Modules
- LLM 지식 베이스 (LLM 이 관리하는 개인 위키)
- MCP — Model Context Protocol
- 기계적 해석가능성
- MHS — Model Hardware Standard
- 군사·정보 역량 평가
- 모델 라우팅 (Model Routing)
- 오픈 웨이트 정책 공방
- 사후 학습 스케일링
- Preparedness Framework
- R&D 자동화 지수 (R&D Automation Index)
- 추론 모델
- 세이프티 케이스 (Safety Cases)
- 안전 모니터링과 데이터 보관
- Software 3.0
- 테스트 타임 연산 (추론 시점 연산 스케일링)
- 월드 모델 (World Models)
조직 · 35
- AI Evaluator Forum (AEF)
- Ai2 (Allen Institute for AI)
- Aleph Alpha
- Alibaba / Qwen AI Lab
- AMD
- Ant Group (inclusionAI / AntLing)
- Anthropic
- Apodex
- Apple
- DeepSeek
- Discovery Loop
- Google DeepMind
- Hugging Face
- IBM
- Institute of Foundation Models (IFM)
- Liquid AI
- Meituan
- Meta AI
- Microsoft
- MiniMax
- Mistral AI
- Moonshot AI
- NVIDIA
- OpenAI
- Poolside
- PrismML
- Runway
- Sakana AI
- Tencent
- Thinking Machines Lab
- TypeSafe AI
- World Labs
- xAI
- Xiaomi
- Z.ai
모델 · 121
- AlphaEvolve
- Astra
- ChatGPT Images 2.5
- Claude Fable 5
- Claude Fable 5.1
- Claude Mythos Preview
- Claude Opus 4.7
- Claude Opus 4.8
- Claude Opus 5
- Claude Opus 5.5
- Claude Science
- Claude Sonnet 5
- Claude Sonnet 5.5
- Co-Scientist (Google DeepMind)
- Cosmos 3 Super
- Cosmos-H-Dreams
- Deep Research Max
- DeepSeek V4
- DeepSeek V4-Flash
- DeepSeek V4-Flash-Vision-Exp
- DeepSeek V4-Pro-0813
- DeepSeek V4.1-Flash
- Devstral 2
- DiffusionGemma
- Fugu Max
- Fugu Ultra v2
- Gemini 3.1 Deep Think
- Gemini 3.5 Flash
- Gemini 3.5 Flash Cyber
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Pro
- Gemini 3.5 Transcribe
- Gemini 3.6 Flash
- Gemini 3.7 Flash
- Gemini 3.8 Flash
- Gemini 3.8 Flash Cyber
- Gemini 3.8 Flash TTS
- Gemini 3.8 Flash-Lite TTS
- Gemini 3.8 Live
- Gemini 3.8 Live Extended Thinking
- Gemini 4
- Gemini 4 Argon
- Gemini Omni
- Gemini Omni 1.1 Flash
- Gemini Robotics 2
- Gemini Robotics ER 1.6
- Gemini Robotics ER 2
- Gemini Spark
- Gemma 3n
- Gemma 4 12B
- GLM-5.2
- GLM-5.3
- GLM-5.3-Flash
- GPT-5.5 Instant
- GPT-5.6 Sol (및 Terra, Luna)
- GPT-5.6-Cyber
- GPT-6 Luna
- GPT-6 Sol
- GPT-6.1 Sol
- GPT-Image-2.5 Flare
- GPT-Image-2.5 Sunburst
- GPT-Live-1
- GPT-Realtime-2 (OpenAI)
- GPT-Rosalind
- Granite 4.2
- Grok 4.1 Fast (xAI)
- Grok 4.5
- Grok 4.6
- Grok 4.7
- Grok Build
- Grok Imagine Image 2.0
- Grok Imagine Video 1.5 (Preview)
- Grok V9-Medium
- Grok Voice Think Fast 2.0
- Grok Voice Transcribe 2.0
- GWM Worlds 2
- Hy4 preview
- Inkling
- Jev
- K2 Horizon
- Kimi K2.8 Preview
- Kimi K3
- Kolibri-1
- Laguna S 2.1
- Leanstral 1.5
- LFM2.5-2.6B
- Ling-3.0-tiny
- LongCat-2.0
- Lyria 3.5
- MAI-Code-1 / MAI-Code-1-Flash
- MAI-Thinking-1
- MiMo-V2.6-Pro
- Ming-Image-0.1-Design
- MiniMax H3
- MiniMax M3
- MiniMax Music 3.0
- Mistral Large 3
- Mistral Medium 3.5
- Muse Glimmer
- Muse Image
- Muse Realtime Avatar
- Muse Spark (1.0 / 1.1)
- Muse Spark 1.2
- Muse Spark 1.3
- Muse Video
- Muse Voice Transcribe
- Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
- Nemotron 3.5 Lightning
- NVIDIA Kumo Tabular
- Project Polaris
- Qwen 3.8 27B
- Qwen 3.8 Max
- Qwen-Drive-1.0-4B
- Qwen-Image-2.1
- Qwen3.8-Flash-Next
- Robostral Navigate
- Shieldstral 1.0
- SL2T
- Ternary Bonsai 2 27B
- WeatherNext 3
- WeatherNext Cyclones