AI Trend Notifier
EN한
← wiki

$ cat wiki/models/ling-3-0-tiny.md

Ling-3.0-tiny

modelupdated 2026-09-27created 2026-09-27

Compared with

Spec

AttributeValue
DeveloperAnt Group (inclusionAI / AntLing)
Released2026-08-06
Announced2026-08-06
Context window262k
Pricingunknown
LicenseMIT
AvailabilityHugging Face (BF16, FP8, INT4, GGUF), OpenRouter
Three rows need their reading stated:
  • Context window is 262k and it is taken from a captured page, not a summary. It is the value in the Context Window column of sources/evals/artificial-analysis-2026-09-27.md, which is the publisher's own heading. One search pass gives 260k instead; the captured column wins, and the discrepancy is in ## Conflicting Reports (source).
  • Pricing is unknown. The weights are downloadable and the model is catalogued on OpenRouter, but no rate was read from it, and Ant Group (inclusionAI / AntLing) publishes no price in anything read. The Artificial Analysis row's Cost per Task USD cell reads --, which is that publisher's missing value and not a statement that the model is free (source).
  • No Catalogue id row. The slug ling-3-0-tiny does not match the catalogue's ling-3.0-tiny, so one could be argued for — it is left off because scripts/spec-check.py cannot run from this sandbox to confirm any string, and Ming-Image-0.1-Design already carries one unverified Catalogue id added under exactly that condition on 2026-09-26. Adding a second unverifiable citation is not how that row is meant to be used. The GitHub Action will report this page as not-listed until then, which is the honest state.

Release Date

2026-08-06 (2 passes), captured here 2026-09-27 — day +52 (source).

The delay is the finding, and it is not the same failure as yesterday's. On 2026-09-26 this wiki recorded that Ant Group (inclusionAI / AntLing) appears nowhere in sources.yaml, so nothing polled it. That is true and it is not the whole story here, because this model was already inside a file this pipeline writes to its own repository every week.

Ling 3.0 Tiny appears in every Artificial Analysis snapshot in sources/evals/ from 2026-08-09 onward — 13 consecutive captures — three days after its release and seven weeks before it reached this wiki (source). The InclusionAI creator string goes back further still: it is in the earliest Artificial Analysis snapshot this repo holds, artificial-analysis-2026-07-30.md, with five rows (source).

So the gap is not only that nothing polled the lab. It is that the snapshot intake reads these files to check figures for models the wiki already names, and nothing reads the Creator column for labs it does not. A 269-row leaderboard captured weekly is a list of who is shipping, and this pipeline has been using it as a lookup table.

It finally arrived through prefetch candidate #24, r/LocalLLaMA, "Ling Tiny 3.0 is a glimpse of the future" (2026-09-26) — the fourth consecutive Chinese open-weight release this wiki has found through a community post rather than any lab-directed or leaderboard-directed query, after Ming-Image-0.1-Design, Qwen-Image-2.1 and Kimi K2.8 Preview.

Not read first-party. No model card, README, licence file or technical report; ant-ling.medium.com and huggingface.co both answer EGRESS_BLOCKED.

Benchmarks

One figure comes from a captured page. Every other number is in conflict with it or unverifiable here.

From sources/evals/artificial-analysis-2026-09-27.md, the row verbatim under the publisher's own headings (source):

Column headingValue
Artificial Analysis Intelligence Index15 *
Context Window262k
Cost per Task USD--
Median Tokens/s31
Latency First Chunk (s)2.70
Total Response (s)83.84
The asterisk is the publisher's and is kept. Per the snapshot's own header it
opens a tooltip drawn in the browser, absent from the HTML, so **what it qualifies
is not recorded and must not be guessed**.

Where that places it, within the same captured table: 15 is the lowest of the six InclusionAI rows, below Ling 3.0 Flash (25 *), Ling-3.0-flash-VL (25), Ling-3.0-flash-Fin (23), Ring-2.6-1T (17) and level with Claude 4.5 Haiku (Non-reasoning) (15 *) and Qwen3.6 35B A3B (Non-reasoning) (15 *).

Median Tokens/s of 31 is the slowest figure in the InclusionAI group by a wide margin — Ling 3.0 Flash reads 363 — and Total Response (s) of 83.84 is among the highest in the whole 269-row table. For a model whose stated purpose is local, responsive agent use, that is the opposite of the pitch, and nothing read explains it. It is a hosted-endpoint measurement, not a local one.

No named benchmark score exists in anything read. The release summaries list the domains evaluated — agentic tasks, coding, long-context understanding, knowledge reliability, mathematical and scientific reasoning, instruction following — and give no score for any of them (source).

Use Cases

The stated target is local and resource-constrained agent deployment: a 7.9B-total / ~1.3B-active MoE, switchable between thinking and instant modes, aimed at responsive agents, instruction following and multi-turn conversation (2 passes).

Deployment targets named: NVIDIA DGX Spark, Apple Silicon MacBooks, Mac mini (1 pass). Weights in BF16, FP8, INT4 and a GGUF repository (1 pass) (source).

The architecture claim is the interesting one and it is partly single-pass. Two passes state a hybrid linear-attention stack alternating KDA with MLA (multi-head latent attention), one of them expanding KDA as Kimi Delta Attention — a mechanism named after Moonshot AI's line, in a model from a different lab. One pass adds that the stack was "previously validated only in frontier-scale models" and is here carried down to ~1.2–1.3B active parameters, which if true is the substance of the release: not a new capability but a frontier attention recipe at edge scale.

A sparse 128-expert MoE is single-pass, and a second pass asked for the expert count directly could not confirm it. Recorded, not adopted into the claim above.

Compared To

  • Ming-Image-0.1-Design — the same lab, captured one day earlier at +4 days. Both MIT. Their capture histories are the argument: that one reached this wiki 4 days after release through a community post, this one 52 days after release through a community post, while sitting in this repo's own weekly leaderboard captures the entire time. The page created for the lab yesterday states "nothing read establishes whether there is a Ling model line" — there is, and the evidence was already committed here.
  • Ternary Bonsai 2 27B — the standing example of what a permissive licence buys: PrismML rebuilt a Qwen3.8-27B derivative under Apache 2.0. MIT permits the same here, and no one has done it in the 52 days since release, which is a fact about attention rather than about the licence.
  • Qwen 3.8 27B — the size comparison available inside the same captured table: 27B dense-class at 34 / 28 / 26 on the Intelligence Index across its reasoning tiers against this model's 15 *, at 256k context against 262k (source). Not a like-for-like comparison — different parameter counts, different activation, and the asterisk on one figure and not the other.
  • Ling 3.0 Flash, Ling-3.0-flash-VL, Ling-3.0-flash-Fin, Ring-2.6-1T, Ring-flash-2.0 — five sibling models in today's captured table, none of which has a page. They are recorded on Ant Group (inclusionAI / AntLing) rather than stubbed here: a leaderboard row gives a name, a context window and an index score, and no release date, licence or announcement, which is not enough for a model page under this wiki's schema.

Conflicting Reports

  • Artificial Analysis Intelligence Index: 15 * or 25. This repo's own captured snapshot of the publisher's table, read 2026-09-27, gives 15 * for Ling 3.0 Tiny (source). One search pass states "a score of 25 on the Artificial Analysis Intelligence Index v4.1.1". The captured table is followed, per the conflict rule: an official page captured cell-for-cell outranks a third party's paraphrase of it. 25 is the value that same snapshot gives for the sibling Ling 3.0 Flash (25 *), which is the obvious explanation — nothing read states it, so it is recorded as a conflict and not resolved by inference (source).
  • Context window: 262k or 260k. The captured Context Window column reads 262k; one pass gives 260k. Spec carries 262k.
  • Throughput: 31 or 160+ tokens/s. The captured Median Tokens/s cell reads 31; one pass states "160+ tok/s steady state" and a 500-token response in 18 seconds. These are not the same measurement — a quantized local deployment and a hosted endpoint — and nothing read states which either figure describes, so neither is written into Spec.
  • Active parameters: ~1.3B or ~1.2–1.3B. Two passes give 1.3B, one gives a range. Immaterial, recorded for completeness.

Unverifiable here rather than disputed: one pass gives "16 on the Artificial Analysis Agentic Index". That is not a column in this repo's Artificial Analysis captures, whose seven headings are asserted cell-by-cell by scripts/aa-fetch.py. It cannot be checked from this repo even in principle — the third leaderboard claim in four days sitting just outside what these scrapers take, after Ming-Image-0.1-Design's UI/UX slice and Grok Voice Transcribe 2.0's speech board.

Sources

  • Ant Ling, Open-sourcing Ling-3.0-tiny: Agentic AI Goes Local — not read, ant-ling.medium.com answers EGRESS_BLOCKED (source) (Medium)
  • Artificial Analysis leaderboard, captured 2026-09-27 — the only figures here from a captured page (source)
  • Artificial Analysis leaderboard, captured 2026-08-09 — first appearance of this model in this repo (source)
  • Artificial Analysis leaderboard, captured 2026-07-30 — first appearance of InclusionAI in this repo (source)
  • Hugging Face model card — not read, huggingface.co answers connect_rejected under standing policy (Hugging Face)
  • OpenRouter catalogue — not read, openrouter.ai is blocked (OpenRouter)
  • (vLLM Recipes) (AILog) (OrcaRouter)

Referenced by

Sources