AI Trend Notifier
EN
← wiki

$ cat wiki/models/muse-spark-1-3.md

Muse Spark 1.3

Compared with

Spec

AttributeValue
DeveloperMeta / Meta Superintelligence Labs (MSL)
Released2026-09-02
Announced2026-09-02
Context window1,000,000
PricingStandard $1.25/M input · $4.25/M output · Contributor ≈$0.10/M input · ≈$0.20/M output
Licenseproprietary
AvailabilityMuse Code, Meta Model API
Input is text, image and video; output is text
(source).

Every price and the context window are unchanged from Muse Spark 1.2, four weeks earlier — including the contributor tier, which stays roughly 10–20× cheaper in exchange for Meta training on your traffic. This is now the second consecutive Muse Spark release at the same sticker, which makes the data-for-price trade a standing feature of the line rather than a launch promotion.

No open-weight release was announced. That is stated explicitly in the coverage as putting Muse Spark on a different track from Meta's Llama line, and it is the second Meta frontier model in a row to ship closed.

License reads proprietary rather than unknown, which differs from the 1.2 page. The distinction is not a licence document anyone read — it is that the absence of an open-weight release is stated here as a positive fact about this model, where on 1.2 it was simply unrecorded.

Release Date

2026-09-02, roughly four weeks after Muse Spark 1.2 (2026-08-05).

Captured one day late, and the reason is worth recording. The release does not appear in the 2026-09-03 run's intake because the Latent Space AINews digest that surfaced it (state/prefetch.json candidate #39) was published 2026-09-03T04:38Z, after that run's ledger was written. ai.meta.com is a scrape source with no feed, so nothing else in the intake would have caught it — the same structural gap Anthropic's Alignment Science blog has.

Benchmarks

Vendor-stated, as reported. artificialanalysis.ai, venturebeat.com and datacamp.com all answer EGRESS_BLOCKED from this run's sandbox, so every row is second-hand (source):

BenchmarkMuse Spark 1.3Comparison as given
Artificial Analysis Intelligence Index62 (max) · 61 (xhigh)behind Claude Fable 5.1 and Claude Opus 5 only
DeepSWE v1.175.455.0 for Muse Spark 1.2 · 74.0 for Claude Opus 5
MRCR (long context)98.5 and 98.1GPT-5.6 Sol (and Terra, Luna) 91.5 / 73.8 · Muse Spark 1.2 66.3 / 55.5
GDPVal-AA v21754GPT-5.6 Sol (and Terra, Luna) 1710 · Claude Opus 5 1824
Meta's engineers are quoted as measuring the model completing coding work with
roughly 20% fewer tool calls and 25% fewer tokens than 1.2.

The DeepSWE jump, 55.0 → 75.4, is the largest single-generation move on that benchmark this wiki holds — and it is also the row where the comparison is least safe to read at face value, for the reason below.

The tier the scorecard was measured in is not the tier that shipped

Two published reasoning variants exist: max (top reasoning) and xhigh (faster). max is not generally available at launch — Meta's blog says max reasoning is "coming shortly after we finish additional safety testing" — so the mode a developer can use on day one is xhigh (source).

One pass states that Meta's scorecard compares 1.3's max against 1.2's xhigh, which makes part of the reported jump a reasoning-tier change rather than a generational one. Nothing read contradicts this, and — the part that matters for the table above — nothing read says which individual rows are max and which are xhigh. The Intelligence Index is the only row that reports both (62 and 61); the other three carry one number each and no tier label.

This page therefore does not treat any single row as a like-for-like generational comparison, and the 55.0 → 75.4 figure is recorded as Meta's own framing rather than adopted as a measurement. Per Eval Harness Configuration, a score whose configuration is not stated is not comparable to one whose is.

Use Cases

Long-running agentic, multi-agent and coding workflows, with the improvements Meta names being context management across extended tasks and coding efficiency — tool calls and tokens rather than pass rates (source). See Agents (LLM Agents).

Compared To

  • Muse Spark 1.2 — the predecessor, four weeks earlier, at identical prices and an identical context window
  • Astra — released the day after, and the only other model with a published DeepSWE v1.1 figure from the same week: 74.1% against 75.4 here. Two vendors' announcements, no shared harness
  • Claude Opus 5 — the comparison Meta chooses on both DeepSWE (74.0) and GDPVal-AA v2 (1824), winning the first and losing the second
  • GPT-5.6 Sol (and Terra, Luna) — the MRCR comparison, and the widest reported margin (98.1 against 73.8)
  • Muse Glimmer, Muse Video, Muse Image — the rest of Meta's Muse line

Referenced by

Sources