arXiv cs.CLSeptember 22, 2026
Beyond the Text: Verifying That Agent-Written Papers Are Backed by Their Artifacts
Excerpt
arXiv:2609.22111v1 Announce Type: new Abstract: Large language model agents are increasingly capable of conducting research autonomously, producing research documents alongside the code and experiments that ostensibly support them. Yet whether the reported findings are consistently supported by corresponding implementations and execution evidence remains largely unexplored: existing review practices primarily assess textual quality and cannot reliably identify inconsistencies such as hard-coded