发表机构
Paderborn University; Huawei Hilbert Research Center (Dresden); Technical University of Munich; Huawei Technologies Ltd.; Shanghai Jiao Tong University(帕德博恩大学; 华为希伯特研究中心(德累斯顿); 慕尼黑工业大学; 华为技术有限公司; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MetaPermit提出基于策略的工具访问控制框架,通过LLM推断元属性解耦语义与执行,实现可扩展、一致且可审计的授权,在基准上显著提升任务完成与IPI攻击鲁棒性。
AI 中文摘要
配备工具自主AI代理的兴起带来了重大安全风险,从意外的工具误用到通过间接提示注入(IPI)攻击进行的对抗性操纵。在实践中,已部署的代理系统(如OpenAI Codex和Claude Code)通过粗粒度权限规则和基于LLM的针对单个提议动作的判断相结合来保护工具调用。然而,这两个组件都有重要局限性:静态策略必须预判可能的用户意图,因此无法扩展到开放式任务;而LLM驱动的授权支持动态决策,但会产生不一致的结果,并且仍然容易受到针对性IPI攻击。为了提供可扩展且更一致的授权,我们提出了MetaPermit,一种基于策略的工具访问控制框架,将语义推断与安全执行解耦。通过分析代理-用户交互,我们推导出一组紧凑的、与任务无关的元属性,这些属性捕获用户意图、执行上下文和提议的工具调用之间的关系。这些元属性使MetaPermit无需枚举用户意图即可授权工具使用。在运行时,LLM为每个提议的工具调用推断元属性值,而固定策略评估这些值以允许或拒绝该调用,使每个决策通过推断值和应用的策略规则可审计。我们在AgentDojo和AgentDyn基准上评估了MetaPermit,涵盖七个任务套件和五种攻击方法,使用两个广泛部署的开权重LLM。结果表明,MetaPermit产生的授权决策比LLM驱动的授权一致性高31%,并且在任务完成方面优于最先进的防御方法CaMeL和IPIGuard,改进幅度高达109%,在IPI攻击鲁棒性方面,没有执行恶意工具调用。
英文摘要
The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a combination of coarse-grained permission rules and LLM-based judgments about individual proposed actions. Both components, however, have important limitations: static policies must anticipate possible user intents and therefore do not scale to open-ended tasks, while LLM-driven authorization supports dynamic decisions but produces inconsistent outcomes and remains vulnerable to targeted IPI attacks. To provide scalable and more consistent authorization, we propose MetaPermit, a policy-based tool access-control framework that decouples semantic inference from security enforcement. By analyzing agent-user interactions, we derive a compact, task-independent set of meta-attributes that capture the relationships among the user's intent, the execution context, and the proposed tool call. These meta-attributes allow MetaPermit to authorize tool use without enumerating user intents. At runtime, an LLM infers the meta-attribute values for each proposed tool call, while a fixed policy evaluates these values to allow or deny the call, making each decision auditable through the inferred values and the applied policy rule. We evaluate MetaPermit on the AgentDojo and AgentDyn benchmarks, across seven task suites and five attack methods, using two widely deployed open-weight LLMs. The results show that MetaPermit produces 31% more consistent authorization decisions than LLM-driven authorization and outperforms the state-of-the-art defenses CaMeL and IPIGuard in both task completion, with improvements of up to 109%, and robustness to IPI attacks, with no malicious tool calls executed.