← Back to all articles
arXiv cs.AIOctober 7, 2026

ArtifactArena: Evaluating Models by What They Build in the Physical World

Excerpt

arXiv:2610.06511v1 Announce Type: cross Abstract: To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical de