arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WebFovea:当模型正确但点击错误——面向真实网站视觉型Web智能体的可靠往返机制

WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites

Jiangang Han

arXiv 2610.03036首次发表:更新:

AI 中文总结

WebFovea通过强化模型与浏览器间harness的四个关键阶段并设置护栏,将真实网站视觉Web智能体的隐藏集分数从31.0提升至57.0,获挑战赛第二名。

AI 中文摘要

我们提出WebFovea,一种基于视觉的Web智能体,在WebRetriever挑战赛2026中以满分100分中的57.0分获得第二名。该挑战在WebRetriever基准(arXiv:2607.06118)的协议III上对智能体进行端到端评估:从真实网站的入口URL开始,智能体必须操作网站自身的界面并返回可验证的答案。一个强大的多模态大语言模型(LLM)对此是必要的,但还不够。模型的决策通过harness(模型与页面之间的代码)到达浏览器。在每一步中,四件事必须正确:模型的回复必须被解析为预期的动作,动作必须在页面上生效,结果必须被准确报告回来,并且模型必须被展示其所需的信息。在真实网站上,我们观察到的许多失败发生在这四个阶段之一,而非模型的推理中。一个坐标空间不匹配导致每次点击落在预期坐标的3/4处;对原生下拉框、iframe内和文本框中的操作静默失败;自生成的聊天模板令牌污染了4.9%的任务片段。WebFovea强化了每个阶段,并在循环周围设置护栏,使智能体保持在规则和预算之内。四阶段视图不依赖于模型,尽管某些个别修复确实依赖。因为我们在所有四次提交中使用了相同的模型,我们官方隐藏集分数从31.0上升到57.0反映了harness的变化,以及真实网站上运行间差异的影响。我们描述了设计、每个组件的证据(包括负面结果)、失败分析、局限性以及包括将不同步骤路由到不同模型的路线图。

英文摘要

We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model's decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model's reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately, and the model must be shown the information it needs. On real websites, many of the failures we observed occurred at one of these four stages rather than in the model's reasoning. A coordinate-space mismatch placed every click at 3/4 of its intended coordinates; actions on native dropdowns, inside iframes, and in text boxes failed silently; and self-generated chat-template tokens contaminated 4.9% of task episodes. WebFovea hardens each stage and surrounds the loop with guardrails that keep the agent within the rules and its budget. The four-stage view does not depend on the model, although some individual fixes do. Because we used the same model in all four submissions, the rise of our official hidden-set score from 31.0 to 57.0 reflects changes to the harness, up to run-to-run variance on live sites. We describe the design, the evidence for each component (including negative results), a failure analysis, the limitations, and a roadmap that includes routing different steps to different models. Code is available at https://github.com/jianganghan/WebFovea.

Comments10 pages, 4 figures, 7 tables. Technical report of the 2nd-place solution in the WebRetriever Challenge 2026. Code: https://github.com/jianganghan/WebFovea. v2: added code link

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑