arXiv cs.LGOctober 2, 2026
MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Excerpt
arXiv:2610.00604v1 Announce Type: new Abstract: Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every ta