$ cat wiki/models/kimi-k3.md
Kimi K3
modelupdated 2026-09-30created 2026-07-20
Compared with
Spec
| Attribute | Value |
|---|---|
| Developer | Moonshot AI |
| Released | 2026-07-16 |
| Announced | 2026-07-16 |
| Context window | 1M tokens |
| Pricing | $0.30/M cache-hit input · $3.00/M cache-miss input · $15.00/M output |
| License | Modified MIT (commercial use permitted; weights publicly downloadable) |
| Availability | Kimi Code, Kimi app, Kimi API; weights on Hugging Face (moonshot-ai/kimi-k3, released 2026-07-27); Amazon Bedrock from 2026-09-18 |
Release Date
July 16, 2026.
Benchmarks
- Frontend Code Arena: beats Anthropic Fable 5 (human-preference Elo) — Moonshot's reported benchmark
- Positioned as strongest open-weight model at launch for agentic coding
Use Cases
- Long-horizon agentic coding (K3 Swarm Max variant)
- Chat and general assistant tasks (K3 Max variant)
- Parallel agent workloads (K3 Swarm Max)
Availability changes after release
- 2026-09-18 — Kimi K3 becomes available on Amazon Bedrock, giving a hosted API
option alongside the weights. Captured 2026-09-30, +12 days, on the routine's
Moonshot rotation slot; the release itself is not new and only the surface
changed. One search pass, no first-party read —
huggingface.coand Moonshot's own hosts are blocked from this pipeline. - 2026-09-08 — Unsloth published an updated K3 usage guide aimed at running the model without a multi-GPU cluster. Noted because it bears on who can actually serve a 2.8T open-weight model, and recorded as a third-party artefact.
No Kimi K4 exists. The rotation check found only a single report — The Information, around 2026-07-29 — that Moonshot is seeking additional Blackwell-class GPU capacity. No name, parameter count, benchmark, price or release date has come from Moonshot, and nothing is recorded here as forthcoming.
Compared To
| Model | Params | Context | Open? |
|---|---|---|---|
| Kimi K3 | ~2.8T MoE | 1M | Yes (Modified MIT, released July 27) |
| Qwen 3.8 Max | ~2.4T MoE | unknown | Planned |
| GLM-5.2 | 744B MoE | 1M | Yes (MIT) |
| Claude Fable 5 | unknown | 200K | No |
| MiniMax M3 | 428B MoE | 1M | Yes |
Architecture
- Kimi Delta Attention (KDA): hybrid linear attention mechanism enabling efficient 1M-token processing
- Attention Residuals: enhances long-context coherence
- 896 experts; ~16 active per token (~1.8% activation ratio)
- Native visual understanding
Variants
- Kimi K3 Max — chat and agent tasks
- Kimi K3 Swarm Max — large-scale parallel agent processing
Sources
- Kimi API docs: quickstart
- Simon Willison analysis: kimi-k3
- Tom's Hardware: largest open-weight
Referenced by
AI-Enabled CyberattacksAnthropicAtria Dawn: The Dawn of Agentic SuperintelligenceClaude Mythos PreviewDeepSeek V4-Pro-0813DiffusionGemmaEval Harness ConfigurationFugu MaxGemini 3.6 FlashGLM-5.3Hy4 previewInklingJuly 2026 — Monthly DigestJune 2026 — Monthly DigestK2 HorizonKimi K2.8 PreviewLaguna S 2.1Liquid AIMilitary and Intelligence Capability EvalsMiMo-V2.6-ProMoonshot AIMuse GlimmerOpen-Weights Policy FightPILOT in the Loop: Live Self-Improvement for Long-Horizon Agents (arXiv:2608.26530)PoolsidePost-Training ScalingQwen 3.8 27BQwen 3.8 MaxSakana AIThinking Machines LabWeekly Synthesis — W30 (July 20–26, 2026)Weekly Synthesis — W31 (July 27 – August 2, 2026)When EOS Tokens Disagree: Understanding Length Inflation in On-Policy DistillationXiaomiZ.ai
Sources
- sources/blogs/moonshot-2026-07-16-kimi-k3.md
- sources/blogs/moonshot-2026-07-27-kimi-k3-weights.md
- https://simonwillison.net/2026/Jul/16/kimi-k3/
- https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3
- https://platform.kimi.ai/docs/guide/kimi-k3-quickstart
- https://huggingface.co/moonshot-ai/kimi-k3