AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2610.09484-mimesis.md

MIMESIS — training agents with learned user behavior

paperupdated 2026-10-09created 2026-10-09

TL;DR

MIMESIS learns user behavior from human conversations and supplies a frozen environment for interactive-agent reinforcement learning. (source)

Authors & Org

Hoang Phan, Dat Huynh, Andrey Zhmoginov, Qi Zeng, Wancen Mu, Yue Cao, Shengjie Bi, Yun He, Changdae Oh and Deren Lei. Affiliations are not stated in the abstract record. (source)

Method

The 9B simulator learns 13 behavioral patterns with reasoning supervision. Agent training freezes the simulator. Coached On-Policy Self-Distillation (CSD) converts its private reasoning and subsequent utterances into coaching notes and token-level supervision. (source)

Results

The authors report 65.7 SOUL-Index, 13.4 points higher behavioral fidelity on RealUserSim and 3.6 points lower Turing distance on SimulatorArena than Claude-Opus-5. Across eight environments, agents trained with MIMESIS outperform those trained with GPT-5.5 under all nine unseen user simulators; CSD adds gains under all nine. (source)

Significance

For Agents (LLM Agents) and Agentic Reinforcement Learning, user behavior becomes a learned training resource rather than an off-the-shelf assistant pretending to be a user. (source)

Open Questions

Does the improvement transfer to real people? The abstract establishes generalization to unseen simulators, not real-user deployment performance or training cost. The full paper was not read. (source)

Cite

Referenced by

Sources