← Back to all articles
arXiv cs.AIAugust 18, 2026

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

Excerpt

arXiv:2608.14712v1 Announce Type: cross Abstract: Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first. Standard tools for comparing attention rows (cosine similarity, Jensen--Shannon divergence, Shannon entropy) therefore hinge on a choice papers rarely report: keep the sink, or drop it and renormalize. This choice can reverse conclusions. On ten pretrained mo