← Back to all articles
arXiv cs.AIOctober 7, 2026

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Excerpt

arXiv:2607.06306v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated growing competence in generating web pages from UI screenshots, which convey both visual structure and cues to application behavior. Yet most screenshot-to-code benchmarks emphasize visual fidelity, while interactive generation benchmarks often supply behavioral specifications or demonstrated transitions. Whether models can infer and realize interactions from static screenshots alone remains