← Back to all articles
arXiv cs.AIAugust 17, 2026

From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability

Excerpt

arXiv:2608.08904v2 Announce Type: replace-cross Abstract: How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. First, the VLA decodes depth worse at every layer, a persistent gap we call the floor. Second, the