arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PACE:面向工具使用型LLM智能体的来源感知能力执行

PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

Fengpeng Li, Qizhou Wang, Yuke Hu, Kemou Li, Jun Liu, Haiwei Wu, Jiantao Zhou, Di Wang

arXiv 2610.01349首次发表:更新:

发表机构

King Abdullah University of Science and Technology; RIKEN Center for Advanced Intelligence Project; University of Macau; National Institute of Informatics; University of Electronic Science and Technology of China(阿卜杜拉国王科技大学; 理化学研究所先进智能项目中心; 澳门大学; 国立信息学研究所; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对工具使用型LLM智能体被污染信息诱导的安全问题,提出来源感知能力执行(PACE),通过路径约束与能力效果验证在调用前拦截,显著降低攻击成功率且实用性损失极小。

AI 中文摘要

工具使用型大语言模型(LLM)智能体将生成的文本转化为真实世界中的副作用,因此被污染的工具元数据、检索到的页面、记忆以及可复用的技能都可能影响下一次调用。在准入前对工件进行审查并不能解决这一问题。一个安全变体和一个泄露变体可能产生相同的准入证据,因此一个健全的闸门无法为其中任何一个放宽该站点的限制。我们精确刻画了这一条件,这留下了部署仍可采取行动的最后边界。我们提出了来源感知能力执行(PACE),它在每次工具调用即将执行之前对其进行调解。路径约束提出了所表示影响路径的一个可执行切割,而能力与效果验证则对照从认证请求中编译的权限来检查模式定义的效果。我们将认证执行契约与评估配置区分开来,后者可以在提议的阻断之后恢复授权调用或应用声明的修复。约束要求最终动作保持认证切割。在八个可执行智能体安全基准测试中,针对三个目标模型家族,评估配置在79个符合条件的攻击列中的62个中实现了严格最低的攻击成功率,并在14个中并列;完整基准的原生实用性相对于未防御的智能体最多损失三分。对1167个配对案例的完整消融研究将大部分安全收益归因于效果验证,将拒绝控制归因于边界适应。针对该防御,一种缩小规模的自适应搜索在30个越权目标上全部失败(0/30成功)。

英文摘要

Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑