arXiv cs.AIAugust 18, 2026
Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures
Excerpt
arXiv:2608.14586v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure. In a full autonomous driving stack, this problem is even more pronounced: legacy vehicle platforms were provisioned for modular pipelines, yet after several planning-related functions are absorbed