AI 中文总结
研究指出智能体安全是上下文问题,当前基于内容的框架有误。通过源授权、任务对齐、行动对齐和数据隔离四个属性实施上下文安全,改变了防御连贯性、评估有效性及可识别的攻击模式。
AI 中文摘要
智能体安全通常被视为关于行动内容的问题,防御关注指令是否恶意,基准测试关注智能体是否执行有害行动。本文认为智能体安全本质上是一个上下文相关的问题,当前基于内容的框架系统性地错误定义了它。以‘删除用户数据’命令为例,内容本身无法区分其是常规管理请求还是攻击生产系统的提示注入,而授权上下文可以区分。文中通过四个属性来实施上下文安全,包括源授权、任务对齐、行动对齐和数据隔离。这种重新定义改变了防御的连贯性、评估的有效性以及能识别的攻击模式。
英文摘要
Agent security is widely treated as a question about action content. Defenses ask whether an instruction looks malicious. Benchmarks ask whether an agent performs a harmful sounding action. \textbf{We argue that agent security is fundamentally a contextual problem, and that the current content based framing systematically misdefines it.} A command to ``delete user data'' might be a routine administrative request or a prompt injection attacking production systems, and the content alone cannot distinguish the two. Authorization context can. Across every injection task in AgentDojo and WASP, the same action is one an authenticated user would plausibly request in a routine workflow, which makes the conflation a structural property of evaluating security through content. We operationalize contextual security through four properties that must hold jointly and be evaluated continuously across the agent's trajectory. Source Authorization asks who issued the command. Task Alignment specifies the agent's authorized objective. Action Alignment evaluates whether each action serves that objective. Data Isolation governs information flows across privilege boundaries. Under this reframing, indirect prompt injection becomes a Source Authorization violation. Snapshot benchmarks are structurally incapable of evaluating Data Isolation. Existing defenses are reorganized around the property they actually approximate. The contextual reframing changes which defenses are coherent, which evaluations measure something useful, and which attack patterns evaluation can see at all.
CommentsICML 2026 Position Paper