$ cat wiki/models/ling-3-0-tiny.md
Ling-3.0-tiny
Compared with
- Claude Opus 5.5
- GPT-6 Sol
- MiMo-V2.6-Pro
- Grok 4.7
- Ternary Bonsai 2 27B
- Fugu Max
- Kimi K2.8 Preview
- DeepSeek V4.1-Flash
- K2 Horizon
- Muse Spark 1.3
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- Laguna S 2.1
- Inkling
- LongCat-2.0
- MiniMax M3
- GPT-6 Luna
- Gemini 3.8 Live
- Fugu Ultra v2
- GPT-Image-2.5 Flare
- GPT-Image-2.5 Sunburst
- Astra
- Gemini 3.8 Flash
- Claude Fable 5.1
- GLM-5.3
- Qwen 3.8 27B
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Grok 4.6
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Kimi K3
- GPT-5.6 Sol
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Claude Opus 4.8
- Gemini 3.5 Flash
- Grok Build
- Claude Opus 4.7
- Muse Spark
Spec
| Attribute | Value |
|---|---|
| Developer | Ant Group (inclusionAI / AntLing) |
| Released | 2026-08-06 |
| Announced | 2026-08-06 |
| Context window | 262k |
| Pricing | unknown |
| License | MIT |
| Availability | Hugging Face (BF16, FP8, INT4, GGUF), OpenRouter |
| Three rows need their reading stated: |
Context windowis262kand it is taken from a captured page, not a summary. It is the value in theContext Windowcolumn ofsources/evals/artificial-analysis-2026-09-27.md, which is the publisher's own heading. One search pass gives 260k instead; the captured column wins, and the discrepancy is in## Conflicting Reports(source).Pricingisunknown. The weights are downloadable and the model is catalogued on OpenRouter, but no rate was read from it, and Ant Group (inclusionAI / AntLing) publishes no price in anything read. The Artificial Analysis row'sCost per Task USDcell reads--, which is that publisher's missing value and not a statement that the model is free (source).- No
Catalogue idrow. The slugling-3-0-tinydoes not match the catalogue'sling-3.0-tiny, so one could be argued for — it is left off becausescripts/spec-check.pycannot run from this sandbox to confirm any string, and Ming-Image-0.1-Design already carries one unverifiedCatalogue idadded under exactly that condition on 2026-09-26. Adding a second unverifiable citation is not how that row is meant to be used. The GitHub Action will report this page asnot-listeduntil then, which is the honest state.
Release Date
2026-08-06 (2 passes), captured here 2026-09-27 — day +52 (source).
The delay is the finding, and it is not the same failure as yesterday's. On
2026-09-26 this wiki recorded that Ant Group (inclusionAI / AntLing) appears nowhere in
sources.yaml, so nothing polled it. That is true and it is not the whole story
here, because this model was already inside a file this pipeline writes to its
own repository every week.
Ling 3.0 Tiny appears in every Artificial Analysis snapshot in sources/evals/
from 2026-08-09 onward — 13 consecutive captures — three days after its release
and seven weeks before it reached this wiki
(source).
The InclusionAI creator string goes back further still: it is in the earliest
Artificial Analysis snapshot this repo holds, artificial-analysis-2026-07-30.md,
with five rows
(source).
So the gap is not only that nothing polled the lab. It is that the snapshot
intake reads these files to check figures for models the wiki already names, and
nothing reads the Creator column for labs it does not. A 269-row leaderboard
captured weekly is a list of who is shipping, and this pipeline has been using it
as a lookup table.
It finally arrived through prefetch candidate #24, r/LocalLLaMA, "Ling Tiny 3.0 is a glimpse of the future" (2026-09-26) — the fourth consecutive Chinese open-weight release this wiki has found through a community post rather than any lab-directed or leaderboard-directed query, after Ming-Image-0.1-Design, Qwen-Image-2.1 and Kimi K2.8 Preview.
Not read first-party. No model card, README, licence file or technical report;
ant-ling.medium.com and huggingface.co both answer EGRESS_BLOCKED.
Benchmarks
One figure comes from a captured page. Every other number is in conflict with it or unverifiable here.
From sources/evals/artificial-analysis-2026-09-27.md, the row verbatim under the
publisher's own headings (source):
| Column heading | Value |
|---|---|
| Artificial Analysis Intelligence Index | 15 * |
| Context Window | 262k |
| Cost per Task USD | -- |
| Median Tokens/s | 31 |
| Latency First Chunk (s) | 2.70 |
| Total Response (s) | 83.84 |
| The asterisk is the publisher's and is kept. Per the snapshot's own header it | |
| opens a tooltip drawn in the browser, absent from the HTML, so **what it qualifies | |
| is not recorded and must not be guessed**. |
Where that places it, within the same captured table: 15 is the lowest of the
six InclusionAI rows, below Ling 3.0 Flash (25 *),
Ling-3.0-flash-VL (25), Ling-3.0-flash-Fin (23), Ring-2.6-1T (17)
and level with Claude 4.5 Haiku (Non-reasoning) (15 *) and
Qwen3.6 35B A3B (Non-reasoning) (15 *).
Median Tokens/s of 31 is the slowest figure in the InclusionAI group by a wide
margin — Ling 3.0 Flash reads 363 — and Total Response (s) of 83.84 is
among the highest in the whole 269-row table. For a model whose stated purpose is
local, responsive agent use, that is the opposite of the pitch, and nothing read
explains it. It is a hosted-endpoint measurement, not a local one.
No named benchmark score exists in anything read. The release summaries list the domains evaluated — agentic tasks, coding, long-context understanding, knowledge reliability, mathematical and scientific reasoning, instruction following — and give no score for any of them (source).
Use Cases
The stated target is local and resource-constrained agent deployment: a 7.9B-total / ~1.3B-active MoE, switchable between thinking and instant modes, aimed at responsive agents, instruction following and multi-turn conversation (2 passes).
Deployment targets named: NVIDIA DGX Spark, Apple Silicon MacBooks, Mac mini (1 pass). Weights in BF16, FP8, INT4 and a GGUF repository (1 pass) (source).
The architecture claim is the interesting one and it is partly single-pass. Two passes state a hybrid linear-attention stack alternating KDA with MLA (multi-head latent attention), one of them expanding KDA as Kimi Delta Attention — a mechanism named after Moonshot AI's line, in a model from a different lab. One pass adds that the stack was "previously validated only in frontier-scale models" and is here carried down to ~1.2–1.3B active parameters, which if true is the substance of the release: not a new capability but a frontier attention recipe at edge scale.
A sparse 128-expert MoE is single-pass, and a second pass asked for the expert count directly could not confirm it. Recorded, not adopted into the claim above.
Compared To
- Ming-Image-0.1-Design — the same lab, captured one day earlier at +4 days. Both MIT. Their capture histories are the argument: that one reached this wiki 4 days after release through a community post, this one 52 days after release through a community post, while sitting in this repo's own weekly leaderboard captures the entire time. The page created for the lab yesterday states "nothing read establishes whether there is a Ling model line" — there is, and the evidence was already committed here.
- Ternary Bonsai 2 27B — the standing example of what a permissive licence buys: PrismML rebuilt a Qwen3.8-27B derivative under Apache 2.0. MIT permits the same here, and no one has done it in the 52 days since release, which is a fact about attention rather than about the licence.
- Qwen 3.8 27B — the size comparison available inside the same captured table: 27B dense-class at 34 / 28 / 26 on the Intelligence Index across its reasoning tiers against this model's 15 *, at 256k context against 262k (source). Not a like-for-like comparison — different parameter counts, different activation, and the asterisk on one figure and not the other.
Ling 3.0 Flash,Ling-3.0-flash-VL,Ling-3.0-flash-Fin,Ring-2.6-1T,Ring-flash-2.0— five sibling models in today's captured table, none of which has a page. They are recorded on Ant Group (inclusionAI / AntLing) rather than stubbed here: a leaderboard row gives a name, a context window and an index score, and no release date, licence or announcement, which is not enough for a model page under this wiki's schema.
Conflicting Reports
- Artificial Analysis Intelligence Index: 15 * or 25. This repo's own captured
snapshot of the publisher's table, read 2026-09-27, gives 15 * for
Ling 3.0 Tiny(source). One search pass states "a score of 25 on the Artificial Analysis Intelligence Index v4.1.1". The captured table is followed, per the conflict rule: an official page captured cell-for-cell outranks a third party's paraphrase of it. 25 is the value that same snapshot gives for the siblingLing 3.0 Flash(25 *), which is the obvious explanation — nothing read states it, so it is recorded as a conflict and not resolved by inference (source). - Context window: 262k or 260k. The captured
Context Windowcolumn reads 262k; one pass gives 260k.Speccarries 262k. - Throughput: 31 or 160+ tokens/s. The captured
Median Tokens/scell reads 31; one pass states "160+ tok/s steady state" and a 500-token response in 18 seconds. These are not the same measurement — a quantized local deployment and a hosted endpoint — and nothing read states which either figure describes, so neither is written intoSpec. - Active parameters: ~1.3B or ~1.2–1.3B. Two passes give 1.3B, one gives a range. Immaterial, recorded for completeness.
Unverifiable here rather than disputed: one pass gives "16 on the Artificial
Analysis Agentic Index". That is not a column in this repo's Artificial
Analysis captures, whose seven headings are asserted cell-by-cell by
scripts/aa-fetch.py. It cannot be checked from this repo even in principle — the
third leaderboard claim in four days sitting just outside what these scrapers
take, after Ming-Image-0.1-Design's UI/UX slice and
Grok Voice Transcribe 2.0's speech board.
Sources
- Ant Ling, Open-sourcing Ling-3.0-tiny: Agentic AI Goes Local — not read,
ant-ling.medium.comanswersEGRESS_BLOCKED(source) (Medium) - Artificial Analysis leaderboard, captured 2026-09-27 — the only figures here from a captured page (source)
- Artificial Analysis leaderboard, captured 2026-08-09 — first appearance of this model in this repo (source)
- Artificial Analysis leaderboard, captured 2026-07-30 — first appearance of
InclusionAIin this repo (source) - Hugging Face model card — not read,
huggingface.coanswersconnect_rejectedunder standing policy (Hugging Face) - OpenRouter catalogue — not read,
openrouter.aiis blocked (OpenRouter) - (vLLM Recipes) (AILog) (OrcaRouter)