← Back to all articles
arXiv cs.LGOctober 2, 2026

Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs

Excerpt

arXiv:2610.00524v1 Announce Type: cross Abstract: Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination.