← Back to all articles
arXiv cs.AIAugust 18, 2026

HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals

Excerpt

arXiv:2608.16353v1 Announce Type: cross Abstract: Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments. These models nonetheless carry linearly separable truthfulness signals in their internal representations. Existing white-box detectors, however, collapse this evidence to isolated components or a single depth, discarding discriminative information distributed across the full forward