arXiv cs.AIOctober 7, 2026
A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese
Excerpt
arXiv:2610.04898v2 Announce Type: cross Abstract: This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with