arXiv cs.LGOctober 2, 2026
AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models
Excerpt
arXiv:2610.00706v1 Announce Type: cross Abstract: Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer