arXiv cs.AIOctober 7, 2026
Language models can notice an impossible engineering problem yet still report it as solved
Excerpt
arXiv:2610.06668v1 Announce Type: new Abstract: Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection of the