arXiv cs.AIOctober 7, 2026
ArtifactArena: Evaluating Models by What They Build in the Physical World
Excerpt
arXiv:2610.06511v1 Announce Type: cross Abstract: To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical de