化解爆炸性提示:理解并防范LLM智能体中的触发式提示注入
Defusing Explosive Prompts: Understanding and Preventing Trigger-Based Prompt Injections in LLM Agents
- Tel Aviv University(特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM智能体,提出爆炸性提示(条件触发式注入)攻击,揭示其高成功率与绕过现有防御的能力,并开发检测器DeFuse,通过摄入时检测条件结构有效防御。
AI中文摘要:
随着LLM应用与外部工具集成,它们越来越多地暴露于间接提示注入(IPI)之下,即对抗性指令被嵌入到检索到的内容中。传统的IPI一触即发:智能体一旦摄取内容,就会执行指令。我们引入了爆炸性提示(explosive prompt),这是一种条件性载荷,在攻击者选择的触发条件满足之前保持休眠状态,实际上是在单条检索内容中植入的无需训练、推理时生效的后门。这种时间上的分离达到了普通IPI无法企及的效果。在几乎完全拒绝直接命令的前沿模型上,将同一目标改写为休眠条件形式,能够驱动针对真实智能体后端的实际状态改变工具执行(配对均值16.5%对2.4%,在专有模型上达到34.2%)。在对九个生产智能体(OpenAI Codex、Google Gemini CLI、Anthropic Claude Code CLI、Cursor CLI、GitHub Copilot、Devin AI CLI、Amazon Kiro CLI、Qwen Code、Google Assistant;每组n=30)的试验中,爆炸性提示在43-83%的案例中成功,而直接命令基线最多只有3%,并且它们能绕过已部署的防御:现成的注入分类器对其校准不当,而一个完全封堵直接命令注入的偏好优化模型仍执行了11.8%的爆炸性提示,且每一次都在触发轮次执行。持久的防御杠杆是在摄取时检测条件结构,一旦检测器在爆炸性提示数据上训练——这是之前任何基准都未提供而我们的生成器提供的——重训练将无防御时的实时工具执行攻击成功率从34.3%降至编码器基线的7.5-8.1%。我们的检测器DeFuse在5%校准误报预算下达到3.0%的漏报率,检测质量在所有测试方法中最佳(AUC 0.9994),延迟降低25倍,但需要长度感知阈值。
英文摘要:
As LLM applications integrate with external tools, they are increasingly exposed to indirect prompt injection (IPI), where adversarial instructions are embedded in retrieved content. Conventional IPIs fire on contact: the moment an agent ingests the content, it carries out the instruction. We introduce the explosive prompt, a conditional payload that stays dormant until an attacker-chosen trigger is met, in effect a training-free, inference-time backdoor planted in a single piece of retrieved content. This temporal separation reaches where ordinary IPI cannot. On frontier models that refuse the bare imperative almost entirely, rephrasing the same goal as a dormant conditional drives real, state-changing tool execution against a live agent backend (a paired mean of 16.5% vs. 2.4% for the imperative, reaching 34.2% on a proprietary model). In trials on nine production agents (OpenAI Codex, Google Gemini CLI, Anthropic Claude Code CLI, Cursor CLI, GitHub Copilot, Devin AI CLI, Amazon Kiro CLI, Qwen Code, Google Assistant; n=30 each), explosive prompts succeed in 43-83% of cases versus at most 3% for an imperative baseline, and they slip past deployed defenses: off-the-shelf injection classifiers are miscalibrated on them, and a preference-optimized model that closes imperative injection entirely still executes 11.8% of explosive prompts, every one at the trigger turn. The durable defensive lever is ingestion-time detection of the conditional structure, once detectors are trained on explosive-prompt data, which no prior benchmark supplied and our generator does. Retraining cuts live tool-execution attack success from an undefended 34.3% to 7.5-8.1% for the encoder baselines. Our detector, DeFuse, reaches 3.0% at a calibrated 5% false-positive budget with the best detection quality of any method tested (AUC 0.9994) and 25x lower latency, though it needs length-aware thresholds.