arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ActGuard:针对LLM智能体中间接提示注入的执行前动作审计

ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents

Bingzheng Wang, Xiaoyan Gu, Wentao Wang, Xingyou Yang, Hongcheng Li, Rong Yin

arXiv 2609.14987首次发表:更新:

发表机构

Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences; School of Cyber Science and Technology, Beihang University; University of Wisconsin–Madison(中国科学院信息工程研究所; 中国科学院大学网络空间安全学院; 北京航空航天大学网络空间安全学院; 威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM智能体面临的间接提示注入攻击,提出执行前动作审计框架ActGuard,通过局部工具先验与对比分析识别恶意偏差,在保持任务效用的同时显著降低攻击成功率。

AI 中文摘要

大型语言模型(LLM)智能体通过工具调用与外部环境交互,但工具输出也可能使其面临间接提示注入(IPI)攻击。现有防御主要依赖于提示加固、内容过滤、预生成计划或权限约束。这些方法在处理复杂任务时往往力不从心,或过度净化外部内容,难以在安全性与实用性之间取得平衡。因此,关键挑战在于保留执行灵活性的同时,精确识别并移除真正诱导不安全动作的恶意内容。为应对这一挑战,我们提出了ActGuard,一个执行前动作审计框架。ActGuard并非判断外部内容本身是否可疑,而是评估其是否导致当前动作偏离局部合理的预期。在每一步中,ActGuard预测即将执行的动作可能使用的工具,并在不约束执行轨迹的情况下构建局部工具先验。在执行前,它将候选动作与该先验进行比较,并进行工具级对比分析和参数级证据定位,以识别工具选择和动作参数中的偏差。随后,验证器检查定位到的证据,仅掩蔽被确认为恶意的片段,并从净化后的上下文中重新生成动作。这种设计保留了合法的规划灵活性,同时最大限度地减少了无差别过滤带来的信息损失。我们在具有挑战性的工具使用智能体基准上评估了ActGuard。结果表明,ActGuard将攻击成功率降低到与最先进防御相当的水平,同时将任务效用维持在接近无攻击设置的水平,实现了良好的安全-效用权衡。我们的代码公开于:此https URL。

英文摘要

Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks. Existing defenses mainly rely on prompt hardening, content filtering, pre-generated plans, or permission constraints. These approaches often struggle with complex tasks or over-sanitize external content, making it difficult to balance security and utility. The key challenge is therefore to preserve execution flexibility while precisely identifying and removing the malicious content that actually induces unsafe actions. To address this challenge, we propose ActGuard, a pre-execution action auditing framework. Rather than judging whether external content is inherently suspicious, ActGuard assesses whether it causes the current action to deviate from a locally reasonable expectation. At each step, ActGuard predicts the tools likely to be used by the upcoming action and constructs a local tool prior without constraining the execution trajectory. Before execution, it compares the candidate action against this prior and performs tool-level contrastive analysis and parameter-level evidence localization to identify deviations in tool selection and action parameters. A verifier then examines the localized evidence, masks only spans confirmed as malicious, and regenerates the action from the sanitized context. This design preserves legitimate planning flexibility while minimizing information loss from indiscriminate filtering. We evaluate ActGuard on challenging benchmarks for tool-using agents. Results show that ActGuard reduces attack success rates to a level comparable to state-of-the-art defenses while maintaining task utility close to the no-attack setting, achieving a favorable security-utility trade-off. Our code is publicly available at: https://github.com/binzhwang/ActGuard.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑