← Back to all articles
arXiv cs.CLSeptember 21, 2026

Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

Excerpt

arXiv:2609.21362v1 Announce Type: new Abstract: Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one con