arXiv cs.CLSeptember 11, 2026
Xiaomi-CocktailASR-1 Technical Report
Excerpt
arXiv:2609.11274v1 Announce Type: cross Abstract: Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target