发表机构
Kent State Univeristy; PayPal AI Labs; University of Houston(肯特州立大学; PayPal人工智能实验室; 休斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对间接提示注入下大语言模型智能体的权限漏洞,提出SkillGuard防御机制,通过能力限制策略有效降低攻击成功率,且无额外模型调用或令牌开销。
AI 中文摘要
大语言模型智能体将外部技能的输出放入其执行上下文,使攻击者控制的数据能够影响后续的特权操作。现有防御措施主要对不受信任的内容进行分类或对提议的操作进行授权,并未直接解决当不受信任的数据进入智能体状态后,其未来权限应如何变化的问题。我们提出SkillGuard,这是一个控制层执行层,将该事件视为污染并限制未来能力,以将产生的状态与部署者定义的禁止状态断开连接。在给定可靠的技能摘要和策略的情况下,SkillGuard通过Skill Impact Graph表示与安全相关的转换,通过可操纵性签名指定对技能参数的可允许控制,并通过内联参考监控器来调解调用。在污染发生后,它使用二元、分数或分数流策略计算加权能力限制,无需辅助语言模型推理。我们在四个AgentDojo套件上对SkillGuard进行评估,使用两个后端LLM:Gemini 2.5 Flash和Llama3.3-70B,对比仅使用LLM的无防御基线以及三个不同系统层的防御措施:Spotlighting、CaMeL和AttriGuard。我们构建了一个组合攻击基准,其中每个攻击结合了单独不足以引发目标违规的观察结果,并在该基准上评估相同的基线。在AgentDojo的工具知识攻击下,SkillGuard消除了两个后端在四个套件中三个套件的攻击成功,并将Slack上的攻击成功率分别降低至4.8%和14.3%。针对组合攻击,它在Llama上的表现优于所有基线,在更高的良性效用下与Gemini上最强的基线表现相当。在相同攻击成功率下,分数流限制比二元限制保留了更多的能力。在两种设置中,SkillGuard均未增加模型调用或令牌开销。
英文摘要
Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed operations. They do not directly address how an agent's future authority should change once untrusted data enters its state. We present SkillGuard, a harness-level enforcement layer that treats this event as contamination and restricts future capabilities to disconnect the resulting state from deployer-defined forbidden states. Given sound skill summaries and policies, SkillGuard represents security-relevant transitions with a Skill Impact Graph, specifies admissible control over skill parameters via steerability signatures, and mediates invocations with an inline reference monitor. Following contamination, it computes weighted capability restrictions using binary, fractional, or fractional-flow strategies without auxiliary language-model inference. We evaluate SkillGuard on four AgentDojo suites with two backend LLMs, Gemini 2.5 Flash and Llama3.3-70B, against an LLM-only No Defense baseline and three defenses at different system layers: Spotlighting, CaMeL, and AttriGuard. We construct a compositional attack benchmark in which each attack combines observations individually insufficient to induce target violation and evaluate the same baselines on it. Under AgentDojo's Tool Knowledge attacks, SkillGuard eliminates attack success on three of four suites for both backends and reduces it to 4.8% and 14.3% on Slack. Against compositional attacks, it outperforms every baseline on Llama and matches the strongest baseline on Gemini at higher benign utility. Fractional-flow restriction preserves substantially more capabilities than binary restriction at the same attack success rate. Across both settings, SkillGuard adds no model calls or token overhead.