← Back to all articles
arXiv cs.CLSeptember 11, 2026

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

Excerpt

arXiv:2609.11851v1 Announce Type: new Abstract: Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Traditional language identification models, designed for monolingual text, are not well suited for token-level language identification in code-mixed settings. We formulate the task as