← Back to all articles
arXiv cs.CLSeptember 24, 2026

Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

Excerpt

arXiv:2609.28344v1 Announce Type: cross Abstract: Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs with fewer than 200M parameters. We introduce a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM. Its architec