← Back to all articles
arXiv cs.LGAugust 17, 2026

Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training

Excerpt

arXiv:2608.14498v1 Announce Type: new Abstract: Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using task feedback, but current on-policy RL runtimes execute rollout, reference scoring, and actor training in strict serial phases. While effective for text-only RL, this phase-granular execution is wasteful for VLMs, where processing dense video inputs an