arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAP:用于评估具备复杂动作与感知能力的跨站点浏览器智能体的可扩展基准

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

Zejun Xu, Taiyi Chen, Jin Li, Yongtong Gu, Qi Cheng, Aixuan Lv, Shuai Zhu, Pengfei Zhu, Kaichen Yang, Boyu Sun, Yixian Yang, Mulong Xie, Xin Liu, Dagang Li, Xiaoteng Ma, Hongru Wang

arXiv 2608.08392首次发表:更新:

发表机构

Macau University of Science and Technology; Tsinghua University; Southeast University; FellouAI; ARGUS Lab; The University of Edinburgh(澳门科技大学; 清华大学; 东南大学; FellouAI; ARGUS实验室; 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员推出可扩展基准CAP,通过分解-重组流程构建420项跨站点网页任务,实验发现当前浏览器智能体在感知密集型交互上存在明显瓶颈,与真实需求差距较大。

AI 中文摘要

大型语言模型正越来越多地被部署为通过浏览器与网页交互的自主智能体。尽管近期的进展得益于评估端到端任务成功率的基准,但这些评估在很大程度上忽略了真实网页浏览中两个基本难点:丰富用户界面上的复杂动作,以及动态渲染内容的视觉感知,尤其是在跨越多个网站的工作流中。我们推出CAP,这是一个用于评估浏览器智能体的可扩展基准,针对的是需要非平凡UI交互和视觉理解的跨站点、类人网页任务。具体而言,我们采用分解-重组流程:首先将每个网站抽象为结构化站点卡片,涵盖面向用户的功能、复杂执行操作和感知需求;随后将这些组件重组为逼真的跨站点工作流。因此,每个任务都基于每个网站上的多个特定操作,支持细粒度诊断。在该框架基础上,我们在严格质量控制下,构建了覆盖108个真实网站、24个领域的420个任务。使用我们可验证的“智能体作为评判者”评估框架,对最先进的浏览器智能体开展实验,结果显示成功率较低,且表明感知密集型交互仍是主要瓶颈,暴露了当前智能体与真实网页浏览需求之间的显著差距。

英文摘要

Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven by benchmarks that evaluate end-to-end task success, these evaluations largely overlook two fundamental sources of difficulty in real web browsing: complex actions over rich user interfaces and visual perception of dynamically rendered content, especially in workflows that span multiple websites. We introduce CAP, a scalable benchmark for evaluating browser agents on cross-site, human-like web tasks that require non-trivial UI interactions and visual understanding. Specifically, we adopt a decomposition-and-recomposition pipeline that first abstracts each website into a structured site card capturing user-facing functions, complex execution operations, and perceptual requirements, and then recomposes these components into realistic cross-site workflows. Each task is therefore grounded in multiple specific operations on each website, enabling fine-grained diagnosis. Built on this framework, we construct 420 tasks across 108 real-world websites and 24 domains under careful quality control. Experiments on state-of-the-art browser agents using our verifiable agent-as-a-judge evaluation framework show low success rates and reveal that perception-heavy interactions remain a major bottleneck, exposing substantial gaps between current agents and real-world web browsing demands.

CommentsAccepted to COLM 2026. Project page: https://warriorxu0302.github.io/CAP-Bench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑