← Back to all articles
arXiv cs.CLSeptember 22, 2026

Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models

Excerpt

arXiv:2609.22452v1 Announce Type: new Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio, as a low-compression modality, requires substantially more embeddings than text to preserve both semantic content and acoustic cues. To address this challenge, we introduce \textbf{Vox-Infinity}, the first benchmark spec