AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2610.11794-memento-3.md

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

paperupdated 2026-10-10created 2026-10-10

TL;DR

A frozen LLM agent revises a natural-language rulebook, compiles it into executable world-model code and verifies revisions against observed transitions. Improvement occurs in external memory and code while the underlying model stays fixed. (source)

Authors & Org

Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang, Rasul Tutunov, Haitham Bou-Ammar, Weilin Luo and Jun Wang. Affiliations are not established by the captured abstract record. Submitted 2026-10-08. (source)

Method

The rulebook stores revisable hypotheses and leaves unknown dynamics underspecified. Prediction errors trigger reflection, rule revision and recompilation. New code is accepted only after an LLM faithfulness judgment and cell-exact replay of observed transitions. A population extension shares interaction evidence among multiple world models to guide exploration. (source)

Results

The single-model agent clears every level in 25 public ARC-AGI-3 games, with mean Relative Human Action Efficiency (RHAE) 100.0, using 44% of the human action count. An Atari Pong controller wins 21:0 in each of three tested episodes with different openings, without further LLM calls. These are author-reported public-game and case-study results; the abstract does not establish a private-test score or total learning cost. (source)

Significance

For Agents (LLM Agents) and World Models, this offers a concrete mechanism for learning from interaction without updating LLM weights. Verification checks observed transitions; it does not establish correctness in every unseen state. (source)

Open Questions

  • How well do the learned rules generalize beyond the observed transitions and public games? (source)
  • What are the total exploration and compilation costs, and how sensitive is acceptance to the LLM judge? The abstract leaves these unresolved. (source)

Cite

arXiv:2610.11794. Evidence scope: title, authors and abstract; the full paper was not read. (source)

Referenced by

Sources