arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentTracer:通过细粒度意图-执行对齐追踪间接提示注入攻击

AgentTracer: Tracing Indirect Prompt Injection Attack through Fine-Grained Intention-Execution Alignment

Zitong Yao, Jiangrong Wu, Yixi Lin, Anrui Huang, Luoyun Zhang, Yuhong Nan

arXiv 2610.09935首次发表:更新:

发表机构

Sun Yat-sen University; The Hong Kong University of Science and Technology(中山大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AgentTracer提出意图感知追踪框架,将间接提示注入视为任务意图漂移,通过构建意图驱动执行图恢复隐式依赖,实现攻击链重建与注入源定位,注入点准确率达94.17%。

AI 中文摘要

大型语言模型(LLM)智能体与外部资源交互以完成复杂的用户任务,这使其面临间接提示注入(IPI)攻击,即恶意指令将智能体重定向至攻击者意图的任务。由于IPI在现实环境中难以防御,事后追踪对于定位注入源和重建攻击链至关重要。然而,现有的追踪方法主要捕获显式的控制流和数据流依赖,忽略了由恶意指令驱动的工具调用之间的隐式关系。这些工具调用可能缺乏显式依赖,并与合法操作交织在一起,使得完整攻击链的重建变得困难。在本文中,我们提出了AgentTracer,一个意图感知的追踪框架,将IPI视为任务意图漂移。AgentTracer恢复工具调用之间的隐式决策依赖,构建意图驱动的执行图,通过任务意图连接分散的工具调用。它将用户请求与操作知识库结合,构建用户意图授权空间,并识别意图漂移的工具调用。从审计中的异常工具调用开始,AgentTracer执行基于目标的剪枝和向后追踪,以重建攻击链并定位注入源和注入点。为了在正常任务的噪声存在下评估AgentTracer,我们将AgentDyn和InjecAgent中成功的IPI攻击构建的执行日志与包含1,800个用户请求和9,000个无IPI背景工具调用的正常执行日志相结合。在端到端实验中,AgentTracer实现了94.17%的注入点准确率和93.56%的路径精确率。对比实验表明,AgentTracer在注入点准确率上比现有方法提高了18%至54%。

英文摘要

Large language model (LLM) agents interact with external resources to complete complex user tasks, exposing them to indirect prompt injection (IPI), where malicious instructions redirect agents toward attacker-intended tasks. Since IPI is difficult to defend against in real-world environments, post-incident tracing is essential for locating the injection source and reconstructing the attack chain. However, existing tracing methods primarily capture explicit control-flow and data-flow dependencies, overlooking the implicit relationships among tool calls driven by the malicious instruction. These tool calls may lack explicit dependencies and be interleaved with legitimate operations, making complete attack-chain reconstruction difficult. In this paper, we present AgentTracer, an intent-aware tracing framework that treats IPI as task intent drift. AgentTracer recovers implicit decision dependencies among tool calls to construct an Intent-Driven Execution Graph that connects dispersed tool calls by task intent. It combines the user request with an operation knowledge base to construct a user intent authorization space and identify intent-drift tool calls. Starting from an anomalous tool call under audit, AgentTracer performs target-based pruning and backward tracing to reconstruct the attack chain and locate the injection source and injection point. To evaluate AgentTracer in the presence of noise from normal tasks, we combine execution logs constructed from successful IPI attacks in AgentDyn and InjecAgent with normal execution logs containing 1,800 user requests and 9,000 background tool calls without IPI. In end-to-end experiments, AgentTracer achieves 94.17 percent injection-point accuracy and 93.56 percent path precision. Comparative experiments show that AgentTracer improves injection-point accuracy over existing methods by 18 to 54 percent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑