arXiv cs.LGOctober 7, 2026
Compact Robot Policies Need Fine-Grained Visual Representations
Excerpt
arXiv:2610.08183v1 Announce Type: cross Abstract: Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental. To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model