发表机构
Northern Arizona University(北亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentDrift是一个分步标注的基准,包含12,536条LLM智能体工具调用轨迹,逐步标记注入点与劫持步骤,用于检测间接提示注入攻击。
AI 中文摘要
LLM智能体通过发出工具调用序列来完成任务,它们读取的每个观察结果都是间接提示注入可能进入的通道。当按顺序读取轨迹时,成功的注入具有特征性形状:良性前缀让位于服务于攻击者而非用户的操作。现有基准衡量此类攻击是否对实时智能体成功,现有防护模型将轨迹作为一个整体进行判断;没有公开语料库逐步标注注入进入轨迹的位置以及它破坏了哪些步骤。我们提出AgentDrift,一个包含五个智能体领域的12,536条合成工具调用轨迹的基准,其中71,024个步骤中的每一个都带有四个标签之一:良性、注入点、劫持或失败注入。该语料库包含4,000条良性、5,536条受攻击、1,500条失败攻击和1,500条硬负例轨迹;受攻击轨迹遵循三种合规模式,其标签字符串符合规定的正则文法。失败攻击携带智能体抵抗的注入,硬负例携带类似攻击的合法内容,因此检测器必须区分尝试与成功、偏离与新颖性。轨迹由单个开放模型在类别特定协议下生成,由封闭词汇结构验证器强制执行,由LLM评判者筛选,并在1,200条轨迹上进行人工审计;我们表明LLM评判者本身被硬负例欺骗。表面特征逻辑回归仅恢复55.4%的攻击(F1 0.647),包括仅8.2%的部分劫持和23.1%的延迟执行,因此近一半的攻击需要建模行为序列。我们测量生成数据中的模板集中度、攻击目标族崩溃和世界身份泄漏,并在CC BY 4.0下发布语料库及其文档。
英文摘要
LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection can enter. A successful injection has a characteristic shape when the trajectory is read in order: a benign prefix gives way to actions that serve the attacker rather than the user. Existing benchmarks measure whether such attacks succeed against live agents, and existing guard models judge a trace as a whole; no public corpus labels, step by step, where an injection enters a trajectory and which steps it corrupts. We present AgentDrift, a benchmark of 12,536 synthetic tool-call trajectories over five agent domains in which every one of the 71,024 steps carries one of four labels: benign, injection point, hijacked, or failed injection. The corpus contains 4,000 benign, 5,536 attacked, 1,500 failed-attack, and 1,500 hard-negative trajectories; attacked trajectories follow three compliance patterns whose label strings obey a stated regular grammar. Failed attacks carry an injection the agent resisted, and hard negatives carry legitimate content that resembles an attack, so a detector must separate attempt from success and deviation from novelty. Trajectories were generated by a single open model under category-specific protocols, enforced by a closed-vocabulary structural validator, screened by an LLM judge, and audited by hand on 1,200 trajectories; we show that the LLM judge was itself fooled by the hard negatives. A surface-feature logistic regression recovers only 55.4% of attacks (F1 0.647), including only 8.2% of partial hijacks and 23.1% of delayed executions, so nearly half of the attacks require modeling the behavioral sequence. We measure template concentration, attack-goal-family collapse, and world-identity leakage in the generated data, and release the corpus with its documentation under CC BY 4.0.
Comments19 pages, 8 figures, 13 tables. Dataset: https://github.com/Asif-0209/AgentDrift