arXiv cs.CLSeptember 21, 2026
MemAudit: Auditing Long-Term Agent Memory via Hidden User-State Recovery
Excerpt
arXiv:2606.24595v2 Announce Type: replace Abstract: Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms. In practice, however, this memory is evaluated mostly through downstream behavior, such as later answers, personalization quality, or task success, which tests that understanding only indirectly and leaves the memory artifact itself largely unaudited. We argue that long-term memory shou