arXiv cs.AIOctober 2, 2026
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
Excerpt
arXiv:2610.01780v1 Announce Type: new Abstract: A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up to 12