$ cat wiki/papers/2026/2610.11794-memento-3.md
Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
TL;DR
A frozen LLM agent revises a natural-language rulebook, compiles it into executable world-model code and verifies revisions against observed transitions. Improvement occurs in external memory and code while the underlying model stays fixed. (source)
Authors & Org
Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang, Rasul Tutunov, Haitham Bou-Ammar, Weilin Luo and Jun Wang. Affiliations are not established by the captured abstract record. Submitted 2026-10-08. (source)
Method
The rulebook stores revisable hypotheses and leaves unknown dynamics underspecified. Prediction errors trigger reflection, rule revision and recompilation. New code is accepted only after an LLM faithfulness judgment and cell-exact replay of observed transitions. A population extension shares interaction evidence among multiple world models to guide exploration. (source)
Results
The single-model agent clears every level in 25 public ARC-AGI-3 games, with mean Relative Human Action Efficiency (RHAE) 100.0, using 44% of the human action count. An Atari Pong controller wins 21:0 in each of three tested episodes with different openings, without further LLM calls. These are author-reported public-game and case-study results; the abstract does not establish a private-test score or total learning cost. (source)
Significance
For Agents (LLM Agents) and World Models, this offers a concrete mechanism for learning from interaction without updating LLM weights. Verification checks observed transitions; it does not establish correctness in every unseen state. (source)
Cite
arXiv:2610.11794. Evidence scope: title, authors and abstract; the full paper was not read. (source)