发表机构
University College London; Holistic AI(伦敦大学学院; Holistic AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对Web智能体的观察模式路由问题,发现路由在智能体最需处因标签不足而不可学,固定合适模式的效果优于多数路由策略,仅稀疏单元有例外。
AI 中文摘要
Web智能体通过文本、像素或两者结合的方式观察浏览器,且该选择通常对所有任务固定不变。我们在VisualWebArena和WebArena的8个站点-模型组合(单元)上测量了6种观察模式,并探究按任务选择观察模式能带来什么收益。这些模式具有互补性:每种模式都能解决其他模式遗漏的任务,且失败方式在结构上存在差异,最优选择会在不同任务集之间反转。显而易见的收益是,为每项任务选择最优模式的“神谕”看起来收益巨大,但这一收益被运行间噪声夸大了:对相同任务重复运行同一模式会改变12%-14%的结果,因此对已掌握的模式再运行一次所获得的收益,几乎与新增一个模式相当。真正有效的是成本边界:仅将没有模式能解决的任务分配给成本最低的模式,在8个单元中,有8个单元的成本降低了9.5%-30.6%,且成功率保持不变。随后我们测试了5种路由策略(选择模式、决定何时为强模式付费、从任务文本中读取的零成本规则、置信度级联以及汇总成本层级),没有一种策略能稳健地胜过简单选择一个合适的固定模式;唯一的例外是我们最稀疏单元中的一个脆弱结果。核心障碍在于,路由监督是由智能体的成功率产生的:智能体越弱,路由器获得的标签越少,而这恰恰是路由最具价值的地方。这一限制属于当前的智能体,而非路由本身。标签供给与路由机会同步上升(各单元间相关性为0.95),因此更强的智能体可以推翻这一结果,我们还报告了重复运行的噪声区间和完整测量协议。
英文摘要
Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across eight site-model combinations (cells) on VisualWebArena and WebArena and ask what choosing per task would buy. The modes are complementary: each solves tasks the others miss, they fail in structurally different ways, and the best choice reverses between task sets. The obvious prize, an oracle that picks a winning mode for every task, looks large but is inflated by run-to-run noise: rerunning the same mode on the same tasks changes 12-14% of outcomes, so a second run of a mode already in hand gains about as much as adding a new one. What survives is a cost bound: sending only the tasks no mode solves to the cheapest mode cuts cost by 9.5-30.6% in 8 of 8 cells at unchanged success. We then test five routing policies (picking the mode, deciding when to spend on the strong mode, a zero-cost rule read off the task text, a confidence cascade, and pooled cost tiers), and none robustly beats simply fixing one well-chosen mode; the one exception is a fragile result in our sparsest cell. The central obstruction is that routing supervision is produced at the agent's success rate: the weaker the agent, the fewer labels a router gets, exactly where routing would be most valuable. This limit belongs to today's agents rather than to routing itself. Label supply and routing opportunity rise together (correlation 0.95 across cells), so a stronger agent can overturn the result, and we report the rerun noise bands and the full measurement protocol.
CommentsPreprint. Under review at the Second Workshop for Research on Agent Language Models (REALM), EMNLP 2026 (non-archival track)