← Back to all articles
arXiv cs.LGOctober 1, 2026

Steer-to-Detect: Probing Hidden Representations for Detection of LLM-Generated Texts

Excerpt

arXiv:2605.12890v2 Announce Type: replace-cross Abstract: The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-written text. While recent studies explore leveraging internal representations of language models to uncover deeper detection signals, these raw features often exhibit substantial overlap between classes, limiting their discriminative power. To address this challenge, we propose Steer-to-Detect (\texttt{S2