← Back to all articles
arXiv cs.CLSeptember 22, 2026

When Cosine Similarity Fails to Reflect Linearly Accessible Structure in Dialogue Models

Excerpt

arXiv:2609.22522v1 Announce Type: new Abstract: Cosine similarity is widely used to analyze transformer representations, implicitly assuming that similarity reflects task-relevant structure. We study when this assumption fails in dialogue-conditioned large language models. Across three 7-8B chat-tuned models, ambient cosine similarity substantially underestimates linearly decodable persona structure on the same hidden states; numerically, linear probe AUC is in the 0.73-0.97 range while cosine k