$ cat wiki/papers/2026/2610.09484-mimesis.md
MIMESIS — training agents with learned user behavior
TL;DR
MIMESIS learns user behavior from human conversations and supplies a frozen environment for interactive-agent reinforcement learning. (source)
Authors & Org
Hoang Phan, Dat Huynh, Andrey Zhmoginov, Qi Zeng, Wancen Mu, Yue Cao, Shengjie Bi, Yun He, Changdae Oh and Deren Lei. Affiliations are not stated in the abstract record. (source)
Method
The 9B simulator learns 13 behavioral patterns with reasoning supervision. Agent training freezes the simulator. Coached On-Policy Self-Distillation (CSD) converts its private reasoning and subsequent utterances into coaching notes and token-level supervision. (source)
Results
The authors report 65.7 SOUL-Index, 13.4 points higher behavioral fidelity on RealUserSim and 3.6 points lower Turing distance on SimulatorArena than Claude-Opus-5. Across eight environments, agents trained with MIMESIS outperform those trained with GPT-5.5 under all nine unseen user simulators; CSD adds gains under all nine. (source)
Significance
For Agents (LLM Agents) and Agentic Reinforcement Learning, user behavior becomes a learned training resource rather than an off-the-shelf assistant pretending to be a user. (source)
Open Questions
Does the improvement transfer to real people? The abstract establishes generalization to unseen simulators, not real-user deployment performance or training cost. The full paper was not read. (source)
Cite
- arXiv:2610.09484, submitted 2026-10-07; revised 2026-10-08.
- Captured abstract record.