← Back to all articles
arXiv cs.AIOctober 7, 2026

GlitchPatch: Repairing Glitch Tokens in Frozen Language Models via Local Retokenization

Excerpt

arXiv:2610.04399v1 Announce Type: cross Abstract: Glitch tokens are anomalous vocabulary entries that can cause large language models (LLMs) to produce outputs inconsistent with their inputs. Existing repair methods require access to model internals, making them impractical for frozen checkpoints. We investigate whether glitch tokens can be repaired outside the model by optimizing the input tokenization. An empirical study on BPE merge-rule deletion reveals that (1)deleting a glitch token's merg