$ cat wiki/papers/2026/2610.01780-realcompanion.md
RealCompanion: understanding people across long conversations
TL;DR
RealCompanion tests when an AI companion needs a person's conversational history. High average retrieval performance can conceal failure on the rare messages that actually need distant context. (source)
Authors & Org
Arman Behnam, Sunglyoung Kim, Liangwei Yang. The captured abstract page does not establish affiliations. Submitted 2026-10-01; the page identifies its current version as v2, revised 2026-10-02. (source)
Method
The dataset contains ten real relationships with an AI companion: 27,218 messages over up to 120 days. It includes conversations, profiles, personas, chat tests and question tests. Labels cite supporting messages; chat labels also include the reasoning used to produce them. This tests memory use on real messages rather than only questions written to require retrieval. (source)
Results
- The current abstract reports that 3.4% of messages depend on earlier conversation, with the relevant message a median of 2,157 messages back. (source)
- Recent-message retrieval finds the needed message 95.9% of the time overall but 2.2% when it is far back. Tested detectors barely beat chance on real messages. (source)
- Calling the same prior messages memories increases references to the past by 10 to 14 percentage points, including when memory is unnecessary. Three agent systems reconstruct personas at F1 0.71 but also infer unsupported traits. (source)
Significance
For Agents (LLM Agents), deciding whether memory is relevant needs its own evaluation. A pooled retrieval score is insufficient evidence of useful personalization; see Eval Harness Configuration. This is an interpretation of the reported subgroup results, not a claim about every memory system. (source)
Open Questions
How well do these findings generalize beyond ten relationships, and can a detector improve memory use without inventing personal information? The abstract establishes neither an answer nor independent replication. (source)
Cite
arXiv:2610.01780. Read scope: the current abstract page and the saved daily abstract, not the full paper. (source)
Conflicting Reports
The daily snapshot describes 2.2% as retrieval success on probes needing memory; the current arXiv abstract describes it as success when the needed message is far back. These denominators are not asserted to be identical. The snapshot also reports a 31-fold cost difference at equal F1, while the current abstract gives F1 0.71 and omits that cost claim. The Results section follows the current author abstract; the snapshot remains unchanged. (snapshot) (current abstract)