arXiv cs.AIOctober 7, 2026
ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual Memory
Excerpt
arXiv:2610.05097v1 Announce Type: cross Abstract: As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along three axes: curation, organization, and access. We find that local evidence benefits from global context, compact latent representations balance accuracy and visual-context cost, and the utili