arXiv cs.AIOctober 7, 2026
EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation
Excerpt
arXiv:2605.23954v3 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples candidate responses reflecting its inference-time behavior, while a frozen copy of the same backbone processes the correspond