发表机构
Google(谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WebRider提出角色条件意图控制器,将网络任务委托策略形式化为意图契约,通过分层架构实现,经RiderBench基准测试,提升实时网络智能体的策略遵守度与模型性能。
AI 中文摘要
将网络任务委托给智能体,不只是提出问题,还需传递一套策略:要验证什么、如何处理不确定性、哪些偏好重要、何时停止。然而,当前的实时网络智能体仅以最终答案进行评估,忽略了定义委托的策略约束,看似合理的最终答案可能掩盖了对该策略的违反。我们的全面实时审计揭示了这一关键差距:一个强大的控制器能完成99.2%的任务,但仅在38.8%的案例中遵守所有策略约束,完成任务不代表忠实执行。WebRider通过将委托策略形式化为意图契约来弥合这一差距——该契约是目标、约束、证据义务、答案形式以及任务局部角色控制的操作记录,即便网页发生变化也必须保持有效。WebRider采用分层架构:顶层控制器维护契约,中间层将意图实现为受保护的可执行动作,工具层通过浏览器、搜索和地图工具执行这些动作。我们的基准测试RiderBench在42个公共网站上的4096个实时网络契约上评估该设计,审计内部契约状态和可见用户体验,以确定部署是否遵守策略且步骤符合角色一致性。此外,受保护的中间界面还可作为高质量训练信号:通过该界面训练的8B动作策略模型,在固定控制器下的表现优于仅使用可执行动作的基线模型。通过将浏览路径作为首要对象,WebRider构建了一个可审计、可由人类判断且可学习的系统,不会将动作实现与最终答案决策混为一谈。
英文摘要
Delegating a web task involves more than asking a question; it requires transferring a policy: what to verify, how to handle uncertainty, which preferences matter, and when to stop. Yet, current live-web agents are evaluated solely on the final answer, ignoring the policy constraints that define the delegation. A plausible final answer can conceal violations of that policy. Our full live audit reveals this critical gap: a strong controller completes 99.2% of tasks but honors all policy constraints in only 38.8% of cases. Finishing does not imply fidelity. WebRider bridges this gap by formalizing the delegated policy as an intent contract---an operational record of goals, constraints, evidence obligations, answer form, and task-local persona controls that must hold even as web pages change. WebRider employs a hierarchical architecture: a top-layer controller maintains the contract, a middle layer realizes intentions as guarded executable actions, and a tool layer executes these actions via browser, search, and maps tools. Our benchmark, RiderBench, evaluates this design on 4,096 live-web contracts across 42 public websites, auditing both the internal contract state and the visible user experience to determine if a rollout preserved its policy and if the steps were persona-consistent. The guarded middle interface also serves as a high-quality training signal; an 8B action-policy model trained through this interface outperforms executable-only baselines under a fixed controller. By making the browsing path a first-class object, WebRider enables a system that is auditable, human-judgeable, and learnable without conflating action realization with final-answer decisions. Dataset URL: hf.co/datasets/WebRider/WebRider.