arXiv cs.LGAugust 18, 2026
RecurrentGPT: Expressive Depth through Recurrent Modulation in Transformers
Excerpt
arXiv:2608.15062v1 Announce Type: cross Abstract: Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce RecurrentGPT, a recurrent dep