← Back to all articles
arXiv cs.AIAugust 18, 2026

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

Excerpt

arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 1