行动之前,改变状态:面向欺骗性界面下Web代理的前瞻性状态干预
Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces
浏览论文内容
中文总结 AI 辅助
针对欺骗性界面导致Web代理产生未授权后果的问题,提出Veer运行时防御机制,通过前瞻性状态干预改变任务相关Web状态,在多个基准上显著提升安全任务完成率并降低暗模式成功率。
中文摘要 AI 辅助
基于LLM的Web代理能够自主完成用户任务,然而欺骗性界面可能引导它们产生与用户利益相冲突的结果。现有防御措施主要通过阻断、引导或重新规划来干预代理行为。我们识别出一种不同的失败模式:由于当前的Web状态,一个任务有效的动作仍可能实现未经授权的后果。这促使我们将与任务相关的Web状态本身视为运行时控制目标。我们引入了Veer,一种代理端运行时防御机制,它将任务规划留给基础代理,并在提议的动作会产生未经授权后果时对Web状态进行干预。在修改实时环境之前,Veer构建一条朝向安全任务相关状态的前瞻性干预轨迹,并通过运行时基础和验证来执行它。在TrickyArena和WebDecept上,Veer在所有三种评估设置中实现了最高的安全任务完成率,在TrickyArena-Single和TrickyArena-Multi上分别超过次优防御15.9和25.0个百分点,同时将WebDecept上的暗模式成功率降至0.3%。这些增益在暗模式类型以及所有12种代理、模型和基准配置中持续存在。消融研究表明,主动状态干预提供了最大的增益,而前瞻性展开和时间证据贡献了额外的改进。这些结果确立了与任务相关的Web状态作为保护Web代理免受欺骗性结果影响的有效运行时控制目标。
英文摘要
LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users' interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid action can still realize an unauthorized consequence because of the current Web state. This motivates treating task-relevant Web state itself as a runtime control target. We introduce Veer, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence. Before modifying the live environment, Veer constructs a prospective intervention trajectory toward a safe task-relevant state and executes it with runtime grounding and verification. Across TrickyArena and WebDecept, Veer achieves the highest safe task completion in all three evaluation settings, exceeding the next-best defense by 15.9 and 25.0 percentage points on TrickyArena-Single and TrickyArena-Multi, respectively, while reducing dark-pattern success on WebDecept to 0.3%. These gains persist across dark-pattern types and all 12 agent, model, and benchmark configurations. Ablations show that active state intervention provides the largest gain, while prospective rollout and temporal evidence contribute additional improvements. These results establish task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.
发表机构
- Singapore Management University(新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。