← Back to all articles
arXiv cs.CLSeptember 10, 2026

Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses

Excerpt

arXiv:2609.09902v1 Announce Type: new Abstract: Reading a transformer's internal states in token space is easy to do and hard to trust: a logit lens on a single hidden state is dominated, at intermediate layers, by the generic tokens the model would predict for almost any input. We read the difference instead. Subtracting two closely matched prompts' hidden states and projecting through the unembedding cancels the shared component and surfaces what separates them, an operation equivalent to read