$ cat wiki/models/muse-realtime-avatar.md
Muse Realtime Avatar
Spec
| Attribute | Value |
|---|---|
| Developer | Meta AI |
| Released | not yet |
| Announced | 2026-09-23 |
| Context window | unknown |
| Pricing | unknown |
| License | unknown |
| Availability | unknown |
| Four rows need their reading stated | |
| (source): |
Releasedreadsnot yet. It was shown at Meta Connect 2026 and nothing read states a rollout date, a region, a surface or a preview programme. This is thenot yetcase the schema describes:Announcedis known andReleasedis not.Context windowisunknownand the gap is a category mismatch, as on GWM Worlds 2: this generates continuous video and audio, and no token budget is stated. Unlike that page, the output dimensions are stated — 448×768 at 25 fps — and they are in## Benchmarksrather than here, because they are not a context window.PricingandAvailabilityareunknown. Muse itself is reported on $20 and $100 tiers, but nothing read attaches the avatar to a tier, and inferring one would be this wiki asserting a price Meta has not published.Licenseisunknown, notproprietary. Every other Meta Muse entry here is closed-weight, but nothing read says so about this model, and a licence carried over from a sibling is a guess.
Release Date
Announced 2026-09-23 at Meta Connect 2026, captured here 2026-09-27 — day +4 (source).
The four days are a hold being released, not a miss. The Connect keynote was
captured on the day in sources/blogs/meta-2026-09-23-connect-muse-charm.md, and
that capture does not contain this model. The name surfaced on 2026-09-26
in a single pass and was written into entities/meta-ai.md under
## Conflicting Reports as recorded and not adopted, per this wiki's rule
about single-pass items absent from the original capture. On 2026-09-27 three
independent passes name it with specifications, so the hold is lifted and the
entry moves into Recent Activity.
That is the rule working as intended, and it is worth stating plainly because the alternative reads identically at the time: had the 09-26 run adopted it, the page would have carried a model with no resolution, no frame rate, no latency and no watermarking statement — and the watermarking is the part that changes what this wiki concluded about it.
No Meta first-party page was read, on this run or any other:
www.meta.com answers EGRESS_BLOCKED, confirmed again 2026-09-27. Every figure
below is second-hand.
Benchmarks
No benchmark. What exists is a specification and one unquantified preference claim.
Stated specification (source):
| Attribute | Stated value | Passes |
|---|---|---|
| Video resolution | 448×768 | 1 |
| Frame rate | 25 fps | 1 |
| End-to-end latency | ~870 ms | 2 |
| Watermarking | all output watermarked as AI, "without adding latency" | 1 |
| The comparison claim carries no number. Meta is stated to say the model was | ||
| preferred over Runway Characters and HeyGen LiveAvatar in head-to-head tests — | ||
| and nothing read gives a **win rate, a sample size, a rater description, a date, | ||
| or which versions of the comparators were used**. It is recorded as a vendor | ||
| preference claim and not as a result. |
It is still worth recording, for one reason: Runway Characters is a product of Runway, a page this wiki created one day earlier, on GWM Worlds 2. Two unrelated captures, four days apart, and the second names the first as the thing to beat.
Not established: how the 870 ms is measured or against what; whether the figure is median or worst-case; what the resolution is at, and whether 448×768 is a fixed output or a default.
Use Cases
Stated as embodiment for a conversational agent: it pairs with Muse Realtime Voice so that audio and video stay in sync while Muse replies in under a second, for as long as you keep chatting (2 passes).
The input is a reference image — stated examples a photograph, a full-body illustration, an animal, a household object — rendered talking, gesturing and shifting posture in continuous video for the length of the conversation (1 pass).
The shipped default is a character rather than a person: "Jolly", cream-coloured with beady black eyes and an upward smile, described as "jolly and completely customizable" (2 passes).
That default is a product decision with a provenance consequence, and it is the one place this page will editorialise: a cartoon default and "a photograph" in the same feature are different risk surfaces, and the watermark is the only thing read that distinguishes them. See Content Provenance (AI output marking).
Compared To
- Gemini 3.8 Live — the direct comparator, and the reason this page changes a standing observation. On 2026-09-26 this wiki recorded that Gemini 3.8 Live carries SynthID on every generated frame while nothing read stated any watermarking for Meta's, and wrote that gap into Meta AI and GWM Worlds 2. Today's capture states all output watermarked as AI, "without adding latency". The gap closes on Meta's side and it does not close on Runway's — GWM Worlds 2 still has no provenance statement of any kind. The mechanisms are not comparable: SynthID is a named, published scheme; Meta's is described only as "watermarked as AI".
- GWM Worlds 2 — Runway's continuous-video model, 720p at 24 fps, arbitrary duration, steered by timestamped events while generating. The contrast is what each is for: that one generates a world, this one generates a face that is listening. Both hold a frame rate and a resolution and neither holds a benchmark.
- Muse Voice Transcribe — the other half of the Muse voice stack, and
the precedent for this page's
unknownrows: that release also arrived with no model id, no availability surface and no context limit, three rows readingunknownfor the same reason.
Sources
- Meta, Everything We Announced at Meta Connect 2026 — not read,
www.meta.comanswersEGRESS_BLOCKED(source) (Meta) - The 2026-09-23 Connect capture, which does not contain this model — cited for that absence (source)
- (TechCrunch) (NBC News) (OrcaRouter) (Marketing4eCommerce)