arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过受控环境干预构建具有挑战性的浏览器使用任务

Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions

Xunjian Yin, Tianchen Guan, Jinao Wang, Weili Cao, Daisy Xinlei Lin, Royce Cheng-Yue, Keagan Long, Kyle Wong, Bhuwan Dhingra, Xiangjun Wang, Shuyan Zhou

arXiv 2609.35814首次发表:更新:

发表机构

Duke University; Amazon(杜克大学; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对浏览器使用智能体,通过受控环境干预将难度变为可编程属性,构建挑战性任务基准,显著降低智能体通过率,揭示信念失败为主要失败模式。

AI 中文摘要

随着浏览器使用智能体的性能提升,基准测试通过收集新任务、新网站和新应用来保持同步,通常使任务变得更长或更新颖。这使得难度更新成本高昂且难以控制:当许多方面同时变化时,不清楚究竟是什么让任务具有挑战性。我们转而从智能体已能解决的任务中构建具有挑战性的实例,将难度转化为环境的可编程属性。BreakingWeb将每个基础任务与一个干预条件配对,该条件保留用户指令、潜在目标和后端成功标准,同时在不同网络栈层改变环境。每次干预都是确定性的、可检测的、可恢复的,并标注了其主要加载的认知原语。该基准包含519个干净/干预任务对,涵盖七个自托管网站和29个干预家族,全部根据结果进行评分。我们评估了六个强浏览器使用智能体、三个仅看截图的仅GUI智能体以及人类。构建是有效的:干预平均使智能体通过率降低22.9%,并推翻了每个智能体干净解决的近一半任务,而人类在首次尝试中损失10.0%,在一次熟悉尝试后损失5.7%。主要的失败是信念失败:六个智能体的失败中有75%以声明成功结束,尽管所需的变化从未发生。我们的代码、数据和环境在此http URL公开可用。

英文摘要

As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. This makes difficulty expensive to refresh and difficult to control: when many aspects change at once, it is unclear what actually makes a task challenging. We instead construct challenging instances from tasks agents already solve, turning difficulty into a programmable property of the environment. BreakingWeb pairs every base task with an intervention condition that preserves the user instruction, latent target, and backend success criterion while changing the environment at different web stack layers. Each intervention is deterministic, detectable, and recoverable, and is annotated with the cognitive primitive it primarily loads. The benchmark contains 519 clean/intervention task pairs across seven self-hosted websites and 29 intervention families, all graded against outcomes. We evaluate six strong browser-use agents, three GUI-only agents that see only screenshots, and humans. The construction is effective: interventions cut agent pass rate by 22.9% on average and overturn nearly half of the tasks each agent solves cleanly, whereas humans lose 10.0% on a first attempt and 5.7% after one familiarisation attempt. The dominant failure is belief failure: 75% of the six agents' failures end with a declared success although the required change never happened. Our code, data and environment are publicly available at www.breakingweb.app.

Comments40 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑