← Back to all articles
arXiv cs.CLSeptember 11, 2026

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Excerpt

arXiv:2609.10758v1 Announce Type: new Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contempor