arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UI2App:可执行Web应用程序生成中视觉交互推理的基准测试

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Grace Man Chen, Litao Guo, Yifan Wu, Yiyu Chen, Yenchi Tseng, Sicheng Liu, Yuyu Luo, Ying-Cong Chen

arXiv 2607.06306首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology(香港科学与技术大学(广州); 香港科学与技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有网页生成基准测试对交互能力评估不足的问题,引入UI2App基准测试,通过设计端到端管道沿四个维度评估工件,实验发现视觉与交互能力不匹配,复杂交互是瓶颈,为相关研究提供了参考。

AI 中文摘要

大语言模型在网页生成方面能力不断增强,但现有文本驱动方法存在问题,图像驱动范式更贴近实际工作流程。然而,当前基准测试主要关注视觉保真度,缺乏对生成工件交互能力的系统评估。为此引入UI2App基准测试,它由327张截图组成,设计了端到端管道沿四个维度评估工件。实验表明视觉重建与交互实现能力不匹配,跨页面状态等复杂交互仍是瓶颈。

英文摘要

Large language models (LLMs) have demonstrated growing competence in generating web pages from UI screenshots, which convey both visual structure and cues to application behavior. Yet most screenshot-to-code benchmarks emphasize visual fidelity, while interactive generation benchmarks often supply behavioral specifications or demonstrated transitions. Whether models can infer and realize interactions from static screenshots alone remains insufficiently evaluated. We introduce UI2App to evaluate interaction inference: inferring and realizing application behavior from static visual cues without added behavioral guidance. UI2App comprises 600 screenshots organized into 95 state-coherent sets for runnable multi-route web applications. Our end-to-end pipeline evaluates each artifact along three dimensions: executability, visual fidelity, and interaction inference. The interaction metric (IIS) assesses functional correctness and state-management complexity, crediting valid implementations rather than requiring a match to a single reference. Experiments on six frontier vision-language models reveal a marked mismatch between visual fidelity and interaction realization: the visual-fidelity leader scores only 8.1 on IIS, ranking fourth, while the IIS leader achieves 4.5 times that score. High-complexity interactions such as cross-route state persistence remain a major bottleneck, with five of the six models scoring at most 3.5 on this dimension. Overall, these results highlight interaction inference as a key challenge in generating functional web applications from static screenshots.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑