AI Trend Notifier
EN
← wiki

$ cat wiki/models/muse-spark-1-2.md

Muse Spark 1.2

Compared with

Spec

AttributeValue
DeveloperMeta / Meta Superintelligence Labs (MSL)
Released2026-08-05
Announced2026-08-05
Context window1,000,000
PricingStandard $1.25/M input · $4.25/M output · $0.15/M cached · Contributor $0.10/M input · $0.20/M output · $0.002/M cached
Licenseunknown
AvailabilityMuse Code (beta, macOS + Linux), Meta Model API
There are two price tiers, and the cheaper one is paid for in training data. The
original capture recorded only the standard tier; four outlets read on 2026-08-10
agree on a contributor tier at $0.10/$0.20 per Mtok — roughly **12× cheaper on
input and 21× cheaper on output** — granted **in exchange for explicit permission to
use your prompts and completions to train future Meta models**, and capped at **60
requests per minute per team against the standard tier's 3,000**
(source).
This is a data-for-price trade rather than a volume discount, and the rate cap means
it is usable only for low-concurrency work.

License stays unknown rather than proprietary. Third-party coverage characterises the weights as closed and the terms as the Meta Model API preview terms, with the contributor tier adding a training-data grant on top (source) — but no source read has quoted or linked the licence text itself, and a characterisation is not a licence. No weights were released (source).

Release Date

2026-08-05, alongside Muse Code, the terminal coding agent it powers (source).

Benchmarks

Meta's own reported figures, against Muse Spark 1.1:

BenchmarkMuse Spark 1.1Muse Spark 1.2Gain
Terminal-Bench 2.176.2%82.9%+6.7
DeepSWE v1.153.0%59.3%+6.3
Meta publishes its methodology, and the caveats are Meta's own: the model was run
inside Meta's own vendor agent product (Muse Code), in **isolated Daytona cloud
sandboxes**, under an internal Meta evaluation framework, scored **pass@1 averaged
over five attempts** across the 89 tasks of the official Terminal-Bench 2.1 release.
Meta states that its agent tools and system prompts **"may not be specifically tuned for
proprietary third-party models"**, i.e. that competitors may not be measured at their
best (source).

In Meta's own comparison charts, Claude tops all three, with Muse Spark 1.2 second on the benchmarks Meta chose to highlight (source).

Independently measured by Artificial Analysis, which reports a 3-point gain on its Intelligence Index concentrated in agentic evaluations:

MetricMuse Spark 1.1Muse Spark 1.2
GDPval-AA v21371 Elo1631 Elo
Terminal-Bench v2.178%80%
τ³-Banking25%27%
This repo holds no local Artificial Analysis snapshot carrying these columns, so the
figures are quoted as reported and have nothing local to check them against
(source).

Use Cases

  • Muse Code — terminal coding agent, beta on macOS and Linux. Generates code and verifies it, and manages several persistent background sub-agents across a session (source).
  • Code generation, complex debugging, codebase understanding, end-to-end developer workflows — Meta's stated improvement areas over 1.1 (source).
  • Direct API use via the Meta Model API, unchanged in standard-tier price and context from 1.1 (source). Whether the contributor tier also existed at 1.1 is not stated by any source read (source).

The persistent background sub-agents are the product claim worth separating from the model claim: an agent that keeps sub-agents alive across a session is a harness design, and the benchmark figures above were produced with that harness in the loop. Meta's methodology says as much.

Compared To

ModelPrice (in/out per Mtok)ContextTerminal-Bench 2.1
Muse Spark 1.2$1.25/$4.25 · $0.10/$0.20 contributor1M82.9% (Meta) · 80% (AA)
Muse Spark 1.1 (Muse Spark (1.0 / 1.1))$1.25/$4.251M76.2% (Meta) · 78% (AA)
Meta held standard-tier price and context window constant across the upgrade, which
puts that much of the release in the model and the harness rather than the commercial
terms (source). The
contributor tier is a commercial move, though, and a sharp one: at $0.10/$0.20 it
undercuts every frontier coding model on this wiki, and it is priced that way because
the buyer pays the remainder in training rights over their own code
(source).

Conflicting Reports

The Terminal-Bench and DeepSWE figures are reported three different ways, and the page body follows Meta's own.

SourceTerminal-Bench 2.1DeepSWE 1.1SWE-Bench Pro
Meta (self-reported)82.9%59.3%not reported
Artificial Analysis (independent)80%not reportednot reported
kingy.ai, attributing to Meta's report80.053.361.5
The third row disagrees with the first on both shared benchmarks while claiming the
same origin, and carries a SWE-Bench Pro figure no other source read has. The second row
is a different measurement, not a contradiction — a different harness should be
expected to produce a different number, and Meta says so itself. Recorded rather than
reconciled (source).

Open Questions

  • Does 1.2 share the Muse Spark 1.1 backbone, or is it a separate training run? Its relationship to Watermelon is unstated.
  • Parameter count, architecture and training compute — none published.
  • Is Muse Code available outside the US? The Meta Model API public preview was US-developer-only at 1.1.
  • What licence governs 1.2, and does it differ from 1.1's?

Referenced by

Sources