← Back to all articles
arXiv cs.LGOctober 2, 2026

Learning Rate Transfer for Hybrid Transformer-SSM Architectures

Excerpt

arXiv:2610.01172v1 Announce Type: new Abstract: We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with growing state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size