← Back to all articles
arXiv cs.LGOctober 7, 2026

World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models

Excerpt

arXiv:2603.04317v2 Announce Type: replace-cross Abstract: A growing literature shows that variables can be linearly decoded from the activations of large language models (LLMs). These range from properties of the world, such as the locations of cities and the lifetimes of historical figures, to emotions and pain. Such findings are often taken as evidence that language models go beyond surface text statistics and form internal models of the world. We show that static word embeddings (fixed, conte