可验证动作卡片:面向安全自主智能体的可信人在环控制
The Verifiable Action Card: Trustworthy Human-in-the-Loop Control for Secure Autonomous Agents
浏览论文内容
中文总结 AI 辅助
针对智能体浏览器中提示注入等安全威胁,提出可验证动作卡片(VAC)架构,通过来源隔离、默认拒绝确认和执行绑定,将攻击成功率从68%-100%降至0%,同时保持78%合法任务完成率。
中文摘要 AI 辅助
智能体浏览器可以在用户已认证的会话下执行安全敏感动作,使得间接提示注入和欺骗性确认界面直接威胁到动作完整性。当批准提示本身可能受到不可信页面内容或模型生成文本的影响时,现有的人机交互(HITL)防护措施不足。我们提出了可验证动作卡片(VAC),一种架构级防御机制,它从待处理浏览器动作的真实状态和可信意图来源重建批准信息,在可信的浏览器界面框架外带渲染该信息,并将批准绑定到分派时重新验证的确切动作上。VAC结合了来源隔离、真实状态动作描述符、默认拒绝确认、来源感知风险门控以及执行绑定。我们在一个完整的智能体浏览器中实现了VAC,并在一个包含24个场景的基准上进行了评估,该基准涵盖了混淆代理攻击、循环中的谎言对话伪造、间接提示注入、自适应动作替换、来源规避以及合法任务。在评估的各个LLM中,无VAC时的攻击成功率在68%到100%之间,而VAC将每个模型的攻击成功率降至0%,合法任务完成率为78%,误拦截率为0%。这些结果表明,将批准基于实际将执行的动作,为防范提示级防御和传统HITL确认无法可靠防止的安全故障提供了架构级保护。
英文摘要
Agentic browsers can execute security-sensitive actions under a user's authenticated session, making indirect prompt injection and deceptive confirmation interfaces a direct threat to action integrity. Existing human-in-the-loop (HITL) safeguards are insufficient when the approval prompt itself can be influenced by untrusted page content or model-generated text. We present the \emph{Verifiable Action Card} (VAC), an architectural defence that reconstructs approval information from the ground-truth pending browser action and trusted intent provenance, renders it out-of-band in the trusted browser chrome, and binds approval to the exact action re-verified at dispatch. VAC combines provenance fencing, a ground-truth action descriptor, default-deny confirmation, provenance-aware risk gating, and execution binding. We implement VAC in a complete agentic browser and evaluate it on a 24-scenario benchmark covering confused-deputy attacks, Lies-in-the-Loop dialog forging, indirect prompt injection, adaptive action substitution, provenance evasion, and legitimate tasks. Across the evaluated LLMs, attack success without VAC ranges from $68\%$ to $100\%$, whereas VAC reduces attack success to $0\%$ on every model, with $78\%$ legitimate-task completion and a $0\%$ false-block rate. These results show that grounding approval in the action that will actually execute provides architectural protection against security failures that prompt-level defences and conventional HITL confirmation cannot reliably prevent.
发表机构
- Sir Syed CASE Institute of Technology(Sir Syed CASE 理工学院)
- University of Engineering and Technology(工程与技术大学)
- National Textile University(国立纺织大学)
- King Fahd University of Petroleum and Minerals (KFUPM)(法赫德国王石油与矿物大学)
机构由 AI 辅助整理,请以论文原文为准。