arXiv cs.CLOctober 7, 2026
How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing
Excerpt
arXiv:2610.08544v1 Announce Type: new Abstract: Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what t