← Back to all articles
arXiv cs.CLAugust 19, 2026

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

Excerpt

arXiv:2606.23671v4 Announce Type: replace Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. We extend the question to safety contexts and examine how reliably a model can recognize that its own prior response was elicited by an adversarial prefill attack. Across ten open-weight instruction-tuned LLMs from 3B to 70B and four safety benchmarks, no model reliably recognizes its own compromised outputs, with models claiming intent on prefi