← Back to all articles
arXiv cs.CLSeptember 11, 2026

Xiaomi-CocktailASR-1 Technical Report

Excerpt

arXiv:2609.11274v1 Announce Type: cross Abstract: Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target