arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过你训练过的测试:重新评估LLM智能体的提示注入检测器

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

Zhuowen Liu

arXiv 2610.03448首次发表:更新:

发表机构

Japan Advanced Institute of Science and Technology (JAIST)(日本北陆先端科学技术大学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究重新评估了LLM智能体的提示注入检测器,发现公共基准得分无法预测其在智能体中的实际表现,排名迁移性差,建议部署评估应使用智能体自身工具输出并审计训练数据。

AI 中文摘要

LLM智能体越来越多地使用小型提示注入检测器来筛选工具输出,团队根据公共基准上的得分来选择检测器。我们探讨这些得分是否能预测检测器在智能体内部的行为。我们重放了两个智能体基准(AgentDojo和tau-bench)的地面真实工具调用,不借助LLM来获得构造上良性的工具输出,通过差分重放标记注入输出,并评估了十五个检测器,包括Meta的Prompt Guard 2和两个任务感知的LLM评判器,在这些输出以及BIPIA基准上进行了评估。检测排名在基准之间的迁移性很差:BIPIA上最好的检测器在1%误报率下仅捕获了AgentDojo注入的2%,而一个捕获了AgentDojo注入72%的检测器在tau-bench上仅捕获了15%。工具输出上的误报率(范围从零到超过90%)在两个智能体基准之间确实具有迁移性。在训练数据公开的情况下,训练输入的形式解释了结果。BIPIA的领先者是在完整的BIPIA输入上训练的,但看到InjecAgent的攻击字符串作为短提示并不能帮助它在工具输出中找到它们;在两个智能体基准上表现最好的检测器与任何基准都没有共享数据,并且是在智能体风格的输入上训练的。旨在为部署提供信息的评估应使用智能体自身的工具输出,在低误报率下报告检测结果,并审计检测器的训练数据。

英文摘要

LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain tool outputs that are benign by construction, label injected outputs by differential replay, and evaluate fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges on these outputs and on the BIPIA benchmark. Detection rankings transfer poorly between benchmarks: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, and a detector that catches 72% of AgentDojo injections catches 15% on tau-bench. False-positive rates on tool outputs, which range from none to over 90%, do transfer between the two agent benchmarks. Where training data is public, the form of the training inputs explains the results. The BIPIA leader was trained on full BIPIA inputs, but having seen InjecAgent's attack strings as short prompts does not help it find them inside tool outputs; the best detector on both agent benchmarks shares no data with any benchmark and was trained on agent-style inputs. Evaluations meant to inform deployment should use the agent's own tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on.

Comments12 pages, 5 figures, 4 tables. Code: https://github.com/lzwhehe/benign-instruction-bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑