arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35026cs.AIcs.CL

WebPageBench:面向Web智能体的事件级验证与受控UI变体生成

WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents

  • DAIMLD
  • HSE University(高等经济大学)
  • NUST MISIS(国家研究型技术大学莫斯科国立钢铁合金学院)

机构由 AI 辅助整理,请以论文原文为准。

Anton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev, Alena Fenogenova

AI总结:

WebPageBench通过事件日志验证Web智能体任务,支持受控UI变体生成,提供152个任务和24个模型-配置对,揭示声明完成与实际成功间的显著差距。

AI中文摘要:

我们提出WebPageBench,一个用于评估Web智能体的开放框架,其中每个任务都通过界面自身的事件日志进行验证。六个去除品牌标识的仪器化模拟站点(一个市场、一个书店、一个杂货服务、铁路售票、酒店搜索和一个文档柜)在用户或智能体操作时发出带参数的键入事件。任务声明其所需的事件,成功通过匹配这些事件来决定,无需评判模型,也无需抓取渲染后的页面。相同的仪器化支持受控UI变体生成:一个配置开关通过单个控件的不同实现重新渲染任务,而提示和成功条件保持完全一致,从而可以在固定任务规范下测量对界面形式的敏感性。WebPageBench发布包含三个组成部分:152个任务,分为65个典型场景和87个控制变体,涵盖浅色/深色UI模式;一个通用运行器,使用六个浏览器/DOM配置和五个仅截图GUI智能体家族进行评估;以及一个包含24个模型-配置对的公共排行榜。在公共152任务排行榜上,智能体声明完成与日志确认之间的差距达到41个百分点(一个配置声明所有任务完成,但仅满足59%的条件)。

英文摘要:

We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emit typed events with parameters as a user or an agent acts. A task declares the events it requires, and success is decided by matching them, with no judge model and no scraping of rendered pages. The same instrumentation supports controlled UI variation: one configuration switch re-renders a task through a different implementation of a single control while the prompt and the success conditions stay completely identical, so sensitivity to interface form can be measured under a fixed task specification. The WebPageBench release consists of three components: 152 tasks, divided into 65 canonical scenarios and 87 control variants across light/dark UI-modes; a common runner evaluated with six browser/DOM harness configurations and five screenshot-only GUI-agent families; and a public leaderboard of 24 model-harness pairs. On the public 152-task leaderboard the gap between what agents declare finished and what the log confirms reaches 41 points (one configuration declares every task finished and satisfies the conditions on 59%).

↑