← Back to all articles
arXiv cs.LGOctober 7, 2026

Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell

Excerpt

arXiv:2610.07782v1 Announce Type: cross Abstract: Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3