Sapien:面向自主AI智能体的有状态策略引擎
Sapien: A Stateful Policy Engine for Autonomous AI Agents
浏览论文内容
中文总结 AI 辅助
Sapien通过有状态正则表达式策略引擎,在保持智能体效用接近无约束水平的同时,有效阻止恶意工具调用,显著提升长时程任务安全性。
中文摘要 AI 辅助
上下文安全防御通过综合任务特定策略并在智能体的工具调用上强制执行,来防止AI智能体采取恶意行动。然而,在多步骤任务中,哪些操作有效往往取决于智能体已经做过和学到的事情。我们提出了Sapien,一个用于强制执行有状态上下文策略的策略引擎。Sapien策略使用扩展了有状态谓词、延迟策略生成和范围化语义检查的正则表达式来指定允许的工具调用序列。我们表明,Sapien在效用上保持在无约束智能体的几个百分点之内。即使智能体被完全劫持,Sapien的策略也能排除AgentDojo上93-95%的攻击和Toolathlon上62-85%的攻击(在长时程任务上是工具允许列表的两倍)。
英文摘要
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).
发表机构
- Google(谷歌)
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。