← Back to all articles
arXiv cs.LGOctober 1, 2026

Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

Excerpt

arXiv:2609.39807v1 Announce Type: cross Abstract: Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag falsehoo