arXiv cs.CLSeptember 21, 2026
Enhancing Audio Reasoning via Semantic Summary Prediction
Excerpt
arXiv:2609.20849v1 Announce Type: new Abstract: Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning sequences shift attention away from the audio input. To address this, we propose SPARE (Semantic Prediction for Audio REasoning), which introduces a register token aligned with the final conclusion using a cosine similarity