arXiv cs.CLSeptember 24, 2026
Causal Tracing of Audio-Text Fusion in Large Audio Language Models
Excerpt
arXiv:2603.13768v2 Announce Type: replace-cross Abstract: Despite the strong performance of large audio language models (LALMs) in various tasks, exactly how and where they integrate acoustic features with textual context remains unclear. We adapt causal tracing to investigate the internal information flow of LALMs during audio comprehension. By conducting layer-wise and token-wise analyses across DeSTA, Qwen, and Voxtral, we evaluate the causal effects of individual hidden states. Layer-wise an