arXiv cs.LGOctober 1, 2026
RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
Excerpt
arXiv:2609.39929v1 Announce Type: new Abstract: Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well wit