arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02617cs.SEcs.AIcs.MA

WebUIProof:使用UI智能体执行框架对WebUI代码生成器进行基准测试

WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness

Yun-Yun Tsai, Yuning Mao, Shiqi Wang, Junfeng Yang, Sinong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

WebUIProof是一个面向执行的WebUI代码生成基准,通过UI智能体执行框架进行可执行交互测试,评估八个商业大语言模型发现交互需求频繁失败,并证明该框架可为紧凑模型提供强化学习训练信号以提升功能完成度。

中文摘要 AI 辅助

在大规模评估WebUI代码生成是困难的:输出可能编译通过且看起来合理,但在用户交互下却失败,而先前的基准测试主要依赖带有静态检查(构建成功、截图)的自由形式提示,这些检查无法捕捉功能正确性。我们提出了WebUIProof,一个面向执行的基准测试,为两类任务家族的WebUI生成提供结构化规范和密集的可执行交互测试:通用WebUI(例如,仪表盘、游戏、交互式工具)和3D交互式模拟(例如,粒子/星系系统、物理动力学)。WebUIProof包含一个UI智能体执行框架,该框架使用迭代的计划-行动-观察循环在无头浏览器中运行可执行交互测试:它定位DOM元素,执行操作,观察产生的UI/DOM变化,并检查指定的断言。我们评估了八个商业大语言模型,并观察到即使页面成功渲染,基于交互的需求也频繁失败,尤其是在3D模拟界面上。最后,我们展示了UI智能体执行框架可以提供结果级别的训练信号。使用从可执行交互测试中导出的强化学习奖励训练紧凑模型(例如,Qwen2.5 14B和MIMO 7B),可以提高功能完成度,同时减少构建失败。

英文摘要

Evaluating WebUI code generation at scale is difficult: outputs may compile and look plausible yet fail under user interaction, and prior benchmarks largely rely on free-form prompts with static checks (build success, screenshots) that miss functional correctness. We introduce WebUIProof, an execution-oriented benchmark that provides structured specifications and dense, executable interaction tests for WebUI generation across two task families: general WebUIs (e.g., dashboards, game, interactive tools) and 3D interactive simulation (e.g., particle/galaxy systems, physics dynamics). WebUIProof includes a UI-agent harness that runs executable interaction tests in a headless browser using an iterative plan--act--observe loop: it locates DOM elements, performs actions, observes resulting UI/DOM changes, and checks the specified assertions. We evaluate across eight commercial LLMs and observe frequent failures on interaction-based requirements even when pages render successfully, especially on 3D simulation interfaces. Finally, we show the UI-agent harness can provide outcome-level training signals. Training compact models (e.g., Qwen2.5 14B and MIMO 7B) with RL rewards derived from executable interaction tests improves functional completion while reducing build failures.

发表机构

  • Columbia University(哥伦比亚大学)
  • Meta SuperIntelligence Labs(Meta 超级智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑