$ cat wiki/models/mistral-large-4.md
Mistral Large 4
Compared with
- EmbeddingGemma 2
- Kolibri-1
- GPT-6.1 Sol
- Claude Sonnet 5.5
- MiMo-V2.6-Pro
- Grok 4.7
- Ternary Bonsai 2 27B
- Fugu Max
- Kimi K2.8 Preview
- DeepSeek V4.1-Flash
- K2 Horizon
- Muse Spark 1.3
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- Ling-3.0-tiny
- Laguna S 2.1
- Inkling
- LongCat-2.0
- MiniMax M3
- Gemini 4 Argon
- Claude Opus 5.5
- GPT-6 Luna
- GPT-6 Sol
- Gemini 3.8 Live
- Fugu Ultra v2
- GPT-Image-2.5 Flare
- GPT-Image-2.5 Sunburst
- Astra
- Gemini 3.8 Flash
- Claude Fable 5.1
- GLM-5.3
- Qwen 3.8 27B
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Grok 4.6
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Kimi K3
- GPT-5.6 Sol
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Claude Opus 4.8
- Gemini 3.5 Flash
- Grok Build
- Claude Opus 4.7
- Muse Spark
Mistral's multimodal API preview, announced on 2026-10-06. The announcement promises weights by the end of October; it does not establish that weights are available now. (source)
Spec
| Attribute | Value |
|---|---|
| Developer | Mistral AI |
| Released | 2026-10-06 (API public preview) |
| Announced | 2026-10-06 |
| Context window | 1M |
| Pricing | $1.36/M input · $4.18/M output (announcement) |
| License | unknown |
| Availability | Mistral Studio API preview; weights promised by end of October 2026 |
Release, pricing and availability follow the announcement; context follows the model documentation, which identifies mistral-large-4. The sources read do not name a weight licence. (announcement) (documentation) |
Release Date
The API preview launched on 2026-10-06. Mistral says red-teaming with vetted partners and state authorities precedes the weight release. Future open weights are a commitment, not a completed release. (source)
Benchmarks
Mistral reports DeepSWE v1.1 61.7%, SWE-Atlas-QnA 59.4%, Terminal-Bench 4 28.3%, and AutomationBench 59.9% across 657 business workflows. These are figures from the vendor's announcement, not an independently reproduced evaluation. (source)
Its cyber comparisons include models that refuse tasks. Interpretation: access policy can affect the ordering, so these figures should not be read as capability measured under identical safeguards. (source)
Use Cases
The announcement targets coding agents, multimodal document work and enterprise security. It describes shared RL environments and verification components spanning chat, scientific tasks and long-horizon tool use. (source)
Compared To
Mistral's blind coding evaluation with Surge AI assigns ML4 Preview 3.74 on a 1–5 scale, versus 4.22 for Claude Opus 5 and 3.60 for GLM-5.3. This is a separate human evaluation, not the coding benchmark aggregate. (source)
Conflicting Reports
The announcement describes 1 trillion total and 49 billion active parameters; the model documentation says 1.05T total, 52B active and a 1.6B vision encoder. Neither source read reconciles these descriptions. Both are retained without choosing a parameter count. (announcement) (documentation)
The documentation displays both $1.36/$4.18 and $0.68/$2.09 input/output pairs, plus $0.14/$0.07 cached-input values, without an unambiguous rate-mode label in the captured text. The Spec table therefore attributes its prices to the announcement; the lower pair is not adopted as a price cut. (documentation)