← Back to all articles
arXiv cs.AIOctober 7, 2026

Language models can notice an impossible engineering problem yet still report it as solved

Excerpt

arXiv:2610.06668v1 Announce Type: new Abstract: Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection of the