← Back to all articles
arXiv cs.LGOctober 1, 2026

Cross-Layer Discrete Concept Discovery for Interpreting Language Models

Excerpt

arXiv:2506.20040v4 Announce Type: replace Abstract: Interpreting language models remains challenging due to the existence of residual stream, which linearly mixes and duplicates features across adjacent layers, causing single-layer analyses to miss this cross-layer structure. Cross-layer sparse autoencoders (SAEs) address layer mixing but operate in continuous space, where concepts split across many neurons without clear boundaries. We introduce Cross-Layer Vector Quantized-Variational Autoencod