arXiv cs.CLSeptember 21, 2026
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
Excerpt
arXiv:2609.15991v2 Announce Type: replace Abstract: Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and H\'ello) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization. We present the Functionalizer, a lossless pre-tokenizer framework that factors orthographic and structural variations into a compositional opcode/operand prefix stream before tokenization: a canonical