arXiv cs.CLSeptember 21, 2026
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
Excerpt
arXiv:2602.16313v2 Announce Type: replace Abstract: Existing evaluations of agents with memory typically assess memorization and action in isolation. One class of benchmarks evaluates memorization by testing recall of past conversations or text but fails to capture how memory is used to guide future decisions. Another class focuses on agents acting in single-session tasks without the need for long-term memory. However, in realistic settings, memorization and action are tightly coupled: agents ac