$ cat wiki/papers/2026/2609.24308-happyworld-bench.md
HappyWorld-Bench
TL;DR
The first instrument this wiki holds that evaluates world models on whether the world stays reliable while an agent acts in it, rather than on how the video looks. Six capabilities (W1–W6) across three independent tracks — video, spatial, embodied — with 1,138 video prompts, 300 spatial scenes and 254 embodied test cases, 14 video models, 9 spatial systems and 8 embodied candidates evaluated, and human A/B Elo from HappyWorld-Arena alongside automated metrics. Spatial models reach at best 70.14% placement accuracy and 73.33% edit execution (source).
Authors & Org
Not stated in the snapshot — no author list, no affiliation, and none of the 31
evaluated systems is named. arxiv.org answers EGRESS_BLOCKED from this run's
sandbox. Recorded as unknown rather than guessed.
Method
The stated premise is that evaluating a world model requires two things, and the field has been doing one: the quality of the worlds generated, and their consistency and responsiveness under exploration, interaction and modification.
Structure:
- A hierarchical capability framework of six world capabilities, W1–W6, from generative construction to unified world modelling. The six are counted and not named in anything read.
- Three independent evaluation tracks: video world models, spatial world models, embodied world models.
| Track | Scale | Systems evaluated |
|---|---|---|
| Video | 1,138 prompts | 14 |
| Spatial | 300 scenes | 9 |
| Embodied | 254 test cases | 8 |
| Two measurement routes, run together: HappyWorld-Arena organises **human A/B | ||
| comparisons** to derive model-level Elo ratings, and **newly designed automated | ||
| metrics capture behavioural correctness**. |
That pairing is the methodological point. Human A/B answers which looks better; behavioural-correctness metrics answer whether the world did the right thing. A world model can win the first and fail the second, and this is the first benchmark here built to show the gap.
Not stated in the snapshot: the names of W1–W6, the automated metrics, the Elo sample sizes, the rater pool, and which systems were evaluated in any track.
Results
| Track | Reported finding |
|---|---|
| Video | reduced consistency during extended rollouts and revisits |
| Spatial | at best 70.14% placement accuracy, 73.33% edit execution |
| Embodied | struggle to preserve state across multi-step actions and to respond precisely to altered action conditions and physical rules |
| Only the spatial track carries numbers, and both are ceilings — at best — so | |
| the rest of the field of 9 is below them. The video and embodied findings are stated | |
| as directions with no figure. |
The conclusion the authors draw is the one worth quoting: world models must be evaluated "not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions."
Significance
This is the instrument World Models said did not exist, and it was already published when that page was written.
That page was created on 2026-09-26, on the simultaneous arrival of GWM Worlds 2 (a generative world model, judged on whether the stream stays coherent under control) and Agent-Editing World Model: Rethinking World Modeling for LLM Agents (a predictive one, judged on whether acting on it raises task success). Its central observation was that the two senses of the term arrived from unrelated publishers, neither citing the other's tradition, and that there was no instrument on which they could be compared.
HappyWorld-Bench is dated 2026-09-21 — five days before that page was written — and it spans both senses: the video track is generative coherence, the embodied track is whether state survives multi-step action. The gap was in this pipeline's reading, not in the field. Recorded plainly, because the alternative is a wiki that credits itself with noticing an absence that was not there.
The 2026-09-26 run also recorded 2609.28654 Training Object Permanence in World
Models (196 upvotes today, the snapshot's highest) and 2609.28466 The Past Frames
the Future as one-off mentions on that page. With this paper that makes three
world-model evaluation or memory papers inside one week, all measuring the same
property — does the world remember what it did — and the strongest numbers in any
of them are the 70.14% / 73.33% ceilings here.
Against GWM Worlds 2 specifically: Runway's model publishes no benchmark of any kind, and its own caveat is that real-time generation "trades fidelity for speed" with no figure on either side. A video track of 1,138 prompts with human Elo and behavioural metrics is exactly where that trade would become a number. Nothing read states that GWM Worlds 2 is among the 14, and it could not be — it is contact-only.
Open Questions
- What are W1–W6? The framework is the paper's organising claim and the six capabilities are never named in anything read.
- Which systems? 31 evaluated across three tracks, none named, so no finding here attaches to any model this wiki tracks.
- What are the video and embodied numbers? Only the spatial track is quantified.
- Do the Elo ratings and the behavioural metrics agree? The benchmark exists because they might not, and nothing read reports the correlation.
- Is "at best 70.14%" one system on both spatial metrics, or two? It changes whether any single system is competent at both placement and editing.
Cite
arXiv 2609.24308 — HappyWorld-Bench, 2026-09-21. HuggingFace Daily Papers, 2026-09-27, 41 upvotes — a popularity signal from that community and not a quality or importance ranking (source).