arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Ajar:衡量智能体防御中的开放权限

Ajar: Measuring Open Privilege in Agent Defenses

Reshabh K Sharma, Linxi Jiang, Shuo Chen, Zhiqiang Lin

arXiv 2609.26900首次发表:更新:

发表机构

University of Washington; The Ohio State University; Microsoft Research(华盛顿大学; 俄亥俄州立大学; 微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体安全防御仅按攻击成功率和实用性评估的不足,提出Ajar方法,通过重用现有基准构建任务不需要的候选调用,直接衡量开放权限,并在AgentDojo上验证了五种防御的差异。

AI 中文摘要

语言模型智能体通过其被赋予的工具来执行操作。它在处理任务时读取的数据可以改变它使用这些工具的方式。因此,越来越多的安全可靠智能体执行技术位于智能体与其工具之间,旨在在该边界上实施访问控制、信息流或隔离。如今,这些技术在围绕间接提示注入构建的智能体安全基准上进行评估。这些基准通过衡量防御在多大程度上降低成功攻击数量的同时保持智能体的实用性来评判防御效果。一种防御仅根据智能体的执行情况进行评判。它可以在两个指标上得分很高,同时却保留着任务不需要的传输、删除或广泛读取权限。Ajar 直接利用现有基准来衡量这种开放权限。它附加到一个已存在的智能体安全基准上,并重用该基准用于评估自身运行的测试任务、工具模式、参考解决方案和目标状态。对于每个良性任务,它构建任务不需要的候选工具调用,因此允许这些调用就意味着留下了开放权限。这些调用在智能体可能采取行动的每个点都呈现给防御系统。我们通过将 Ajar 附加到 AgentDojo 上来评估它,在 AgentDojo 中,开放权限成为现有攻击成功率和良性实用性之外的第三个维度。我们在五种防御上运行它:Progent、CaMeL、AC4A、Permission Assistant 和 Claude Code 的 Auto 模式。我们观察到它们留下了数量差异很大的开放权限。两种防御泄露的权限量几乎相同,但它们在完成的良性任务上差异很大,而一种防御通过拒绝其任务有权进行的调用来换取部分紧密性。这种开放权限无法从测得的攻击成功率或良性实用性中推导出来。Ajar 的源代码可在 https://github.com/reSHARMA/Ajar 获取。

英文摘要

A language model agent acts through the tools it is given. The data it reads while working on a task can redirect what it does with those tools. A growing set of techniques for safe and secure agent execution therefore sits between the agent and its tools, aiming to enforce access control, information flow or isolation at that boundary. Today these techniques are evaluated on agent-security benchmarks built around indirect prompt injection. Those benchmarks judge a defense by how far it brings the number of successful attacks down while preserving the agent's utility. A defense is judged only on the agent's execution. It can score well on both metrics while holding open a transfer, a deletion or a broad read that no task needed. Ajar measures that open privilege directly using the existing benchmarks. It attaches to an agent-security benchmark that already exists and reuses the tasks, tool schemas, reference solutions and goal states that benchmark uses to grade its own runs. For each benign task it builds candidate tool calls the task does not need, so allowing one is privilege left open. These calls are presented to the defense at every point where the agent could act. We evaluate Ajar by attaching it to AgentDojo, where open privilege becomes a third axis beside the existing attack success and benign utility. We run it on five defenses: Progent, CaMeL, AC4A, Permission Assistant, and Claude Code's Auto mode. We observed that they leave widely different amounts of privilege open. Two defenses leak by almost the same amount yet differ widely in the benign tasks they finish, and one defense buys part of its tightness by refusing calls its tasks were entitled to make. This open privilege cannot be derived from the measured attack success or benign utility. The source code of Ajar is available at https://github.com/reSHARMA/Ajar.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑