发表机构
Mohamed bin Zayed University of Artificial Intelligence; Amazon; Massachusetts Institute of Technology(穆罕默德·本·扎耶德人工智能大学; 亚马逊; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AdvSim2Real通过共同演化任务课程、注入对手和智能体,在冻结网络世界模型中训练,使4B智能体在未见攻击下完成率提升33.6%,并增强鲁棒性与真实浏览器迁移能力。
AI 中文摘要
网络智能体通过阅读和操作第三方编写的页面来完成用户请求,因此页面上植入的指令可能将智能体从用户的目标引开。智能体不能简单地忽略页面,因为页面也包含任务所需的数值和控制。当前的防御方法是在训练前对固定的注入进行微调,而适应训练后模型的攻击者可以绕过这些防御。对抗训练允许攻击者适应,但保持任务固定,因此一旦智能体解决了任务,任务就不再提供学习信号。我们提出了AdvSim2Real,它在冻结的网络世界模型中共同演化任务课程、注入对手和智能体。课程奖励智能体大约一半时间能解决的任务,对手仅奖励成功翻转,即一种将判定成功转为失败的注入。在模拟器中训练使一个4B智能体既更有能力也更具鲁棒性:其完成率在有攻击和无攻击情况下均提升,能抵御从未训练过的前沿模型对手,且能力提升可迁移到真实浏览器。在150个网络任务上,AdvSim2Real在未见对手下的完成率相对基础智能体提升了33.6%。
英文摘要
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.
CommentsCode at https://github.com/Sarim-MBZUAI/advsim2real