← Back to all articles
arXiv cs.LGOctober 7, 2026

SkillFormer: Skill-Decomposed Adaptation for Audio Language Models

Excerpt

arXiv:2610.07533v1 Announce Type: cross Abstract: Audio language models must handle dozens of distinct skills, from pitch comparison and speaker counting to musical tempo estimation and emotion recognition. Joint training on all skills at once causes interference: gains on one skill often come at the cost of another. We propose \textbf{SkillFormer}, which decomposes audio understanding into skill-specific low-rank adapters and composes them at inference time through a learned router. The router