← Back to all articles
arXiv cs.CLSeptember 24, 2026

Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark

Excerpt

arXiv:2609.28090v1 Announce Type: new Abstract: Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail. A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix relevance.