arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DriftNet:用于检测和定位LLM智能体中提示注入的双头轨迹Transformer

DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents

Asif Pinjari, Mithun Paul Saint-Germain

arXiv 2609.10892首次发表:更新:

发表机构

Northern Arizona University(北亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DriftNet提出双头轨迹Transformer,在一次前向传播中同时完成LLM智能体轨迹的入侵分类与逐步骤注入定位,在AgentDrift基准上实现0.983的F1和98.7%的注入点恢复率。

AI 中文摘要

当间接提示注入成功攻击LLM智能体时,损害会在智能体自身行为中显现:一段良性的工具调用前缀、一个被投毒的观察结果,以及一段服务于攻击者的动作后缀。操作者需要知道三个事实:攻击从哪里进入、破坏了哪些步骤、以及明显的投毒是否被抵抗。现有系统要么返回整个轨迹的判定,要么返回单个不安全索引。我们提出DriftNet,一种双头轨迹Transformer,它读取记录的工具调用轨迹并在一次前向传播中回答所有三个问题:一个头将轨迹分类为是否被入侵,第二个头为每个步骤分配四个标签之一(良性、注入点、被劫持、注入失败)。据我们所知,它是第一个产生这种联合输出的监督检测器。一个冻结的句子编码器和四个无身份世界特征嵌入每个步骤;训练好的主干网络参数低于两百万,并通过两个头上的类加权联合目标进行优化,无需访问智能体的模型。在AgentDrift基准的任务不相交划分(12,536条轨迹,71,024个标记步骤)上,通过20配置扫描将超参数敏感性限制在0.011 F1,且测试部分仅评估一次,DriftNet达到轨迹级F1为0.983,在98.7%的被攻击轨迹上精确恢复注入点,被劫持跨度IoU为0.979,在218次被抵抗攻击上零误报,在困难负样本上误报率为2.9%。在相同划分上重新训练的表面基线恢复了11.1%的部分劫持和17.1%的延迟执行;DriftNet达到98.6%和93.2%,同时降低了每个误报率。阅读全部26个残余错误表明,大多数遗漏可追溯到其标记注入观察不携带可读指令的轨迹,我们报告了基准测量的世界身份规律性以及结果。

英文摘要

When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the agent's own behavior: a benign prefix of tool calls, a poisoned observation, and a suffix of actions that serve the attacker. An operator needs three facts: where the attack entered, which steps it corrupted, and whether apparent poison was resisted. Existing systems return either a whole-trace verdict or a single unsafe index. We present DriftNet, a dual-head trajectory Transformer that reads a logged tool-call trajectory and answers all three questions in one forward pass: one head classifies the trajectory as compromised or not, and a second assigns every step one of four labels (benign, injection point, hijacked, failed injection). To our knowledge it is the first supervised detector to produce this joint output. A frozen sentence encoder and four identity-free world features embed each step; the trained trunk, under two million parameters and optimized with a class-weighted joint objective over both heads, needs no access to the agent's model. On the task-disjoint split of the AgentDrift benchmark (12,536 trajectories, 71,024 labeled steps), with a 20-configuration sweep bounding hyperparameter sensitivity to 0.011 F1 and the test part evaluated exactly once, DriftNet reaches trajectory-level F1 of 0.983, exact injection-point recovery on 98.7% of attacked trajectories, hijacked-span IoU of 0.979, zero flags on 218 resisted attacks, and 2.9% flags on hard negatives. A surface baseline retrained on the identical split recovers 11.1% of partial hijacks and 17.1% of delayed executions; DriftNet reaches 98.6% and 93.2% while lowering every false-alarm rate. Reading all 26 residual errors shows that most misses trace to trajectories whose labeled injection observation carries no legible instruction, and we report the benchmark's measured world-identity regularity alongside the results.

Comments15 pages, 8 figures, 11 tables. Companion detector paper to the AgentDrift benchmark (arXiv:2609.06972); dataset at https://github.com/Asif-0209/AgentDrift

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑