发表机构
Independent Researcher
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体间接提示注入评估中的四类静默缺陷,提出可验证的测试平台,修正评分扭曲并推翻错误的能力壁垒结论。
AI 中文摘要
调用特权工具的LLM智能体容易受到间接提示注入(IPI)的攻击,在这种攻击中,嵌入在检索数据中的对抗性指令会劫持智能体的行为。越来越多的研究工作评估针对IPI的防御措施,但这些评估的有效性却很少被审视。我们审计了一个IPI基准测试及其测试平台,并识别出四类缺陷——静默载荷未投递、攻击成功仅依据工具身份而非参数进行评分、误拒率与模型能力不足相混淆,以及缺乏审计追踪——每一类缺陷都会产生一个看似合理、可发表但实际错误的数字。我们通过在缺陷定义和修正定义下对相同的执行轨迹重新评分来量化这种扭曲:在真实智能体行为上,工具身份评分器报告的攻击成功率为21.7%,而真实的参数级成功率仅为1.2%。在最极端的情况下,一个先前报告为62.8%攻击成功率的开放模型在修正后的测试平台下记录为0%——之前的数字很大程度上是未投递载荷和身份级评分的产物。我们发布了一个测试平台,其构造使得每类缺陷都无法表现——机器可检查的载荷放置、参数级攻击者谓词、每个场景的独立环境,以及强制性的轨迹持久化——并用它来报告该领域尚未提供的三个量化指标:受损智能体是否披露攻击、LLM评判防御的完整安全/效用操作曲线,以及将工具调用能力与防御性过度拦截区分开来。修正后的测试平台进一步推翻了一个已报告的“能力壁垒”:一个被认为无法使用工具的模型实际上完全具备该能力,其先前的结果是环境不匹配的产物。我们认为,评估有效性是智能体安全中防御主张的前提条件,而非附注,并提供了一个强制执行这一点的工具。
英文摘要
LLM agents that invoke privileged tools are vulnerable to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack the agent's actions. A growing body of work evaluates defenses against IPI, but the validity of that evaluation is rarely examined. We audit an IPI benchmark and its harness and identify four defect classes -- silent payload non-delivery, attack success scored by tool identity rather than arguments, false-rejection rate conflated with model incapacity, and the absence of an audit trail -- each of which yields a plausible, publishable, and incorrect number. We quantify the distortion by re-scoring identical execution traces under the defective and corrected definitions: on real agent behaviour, the tool-identity scorer reports a 21.7% attack-success rate where the true argument-level rate is 1.2%. In the sharpest case, an open model previously reported at 62.8% registers 0% under the corrected harness -- the prior figure largely an artifact of undelivered payloads and identity-level scoring. We release a harness whose construction makes each defect unrepresentable -- machine-checkable payload placement, argument-level attacker predicates, per-scenario environments, and mandatory trace persistence -- and use it to report three quantities the field does not: whether a compromised agent discloses the attack, the full security/utility operating curve of an LLM-judge defense, and tool-calling capability disentangled from defensive over-blocking. A corrected harness further overturns a reported "capability barrier": a model deemed incapable of tool use is in fact fully capable, its earlier result an artifact of environment mismatch. We argue that evaluation validity is a prerequisite for, not a footnote to, defense claims in agentic security, and provide an instrument that enforces it.
Comments7 pages, 4 figures, 3 tables