← Back to all articles
arXiv cs.LGOctober 2, 2026

AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models

Excerpt

arXiv:2610.00706v1 Announce Type: cross Abstract: Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer