arXiv cs.AIOctober 7, 2026
DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling
Excerpt
arXiv:2610.04933v1 Announce Type: cross Abstract: Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states eq