arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36570cs.CRcs.AI

CounterSteer:通过激活引导抑制间接提示注入

CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

  • Microsoft Azure(微软Azure)

机构由 AI 辅助整理,请以论文原文为准。

Mark Russinovich

AI总结:

CounterSteer通过推理时激活引导抑制间接提示注入,无需微调,在多个模型上显著降低攻击成功率,同时保持高良性效用。

AI中文摘要:

间接提示注入会使LLM智能体将不可信检索文本视为指令。我们提出CounterSteer,一种推理时防御方法,在模型内部抑制该行为。对于每个模型,一个五步流程从成对片段(仅在是否遵循嵌入指令上有所不同)中拟合残差流方向,并且仅当该方向通过预定义的因果和能力门控时才保留它。在部署时,在预填充阶段从每个工具结果令牌中减去该方向。该编辑始终开启——不存在可逃避的检测决策——并且不需要微调、辅助模型或额外令牌,仅需要白盒服务和工具结果跨度边界。在五个开放权重模型(8B-106B,五个供应商谱系)上,防御后的留出攻击成功率从无防御的0.21-1.00降至0.00-0.17,AgentDojo妥协率从0.10-0.49降至0.006-0.079,同时保持93-100%的排版归一化良性效用,在推理受引导内容时存在较大的任务相关成本。一个基准级别的自适应攻击者达到无防御的0.67-0.73,在评估最深入的两个模型上被限制为其约四分之一。在我们测量的能力较强的模型上的防御中,实现较低妥协率的防御要么损失了22-89%的良性效用,要么对服务权重进行了微调。通过部署向量进行的白盒梯度攻击最多妥协52个片段中的2个,并且2,052个重放的人类红队攻击中没有一个成功。CounterSteer在很大程度上中和了指令性接管:黑盒框架搜索破解了18个开发样本中的3个。参数操纵——在原本合法的调用中由攻击者选择的参数——仅被部分抵抗(18个中的13个);该决策在参数发出时变得线性可读,但在所检查的预生成位置不可读,并且未被所测试的预填充或解码时引导移除,这促使需要参数来源控制。

英文摘要:

Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages), held-out attack success falls from 0.21-1.00 undefended to 0.00-0.17 defended, and AgentDojo compromise rate from 0.10-0.49 to 0.006-0.079, at 93-100% typography-normalized benign utility, with larger task-dependent costs when reasoning over steered content. A benchmark-level adaptive attacker reaching 0.67-0.73 undefended is held to roughly a quarter of that on the two most deeply evaluated models. Among the defenses we measured on capable models, those achieving lower compromise rates either lost 22-89% of benign utility or fine-tuned the served weights. White-box gradient attacks through the deployed vector compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeeds. CounterSteer largely neutralizes instructional takeover: a black-box framing search cracks 3 of 18 development samples. Parameter manipulation--attacker-chosen arguments in otherwise legitimate calls--is only partially resisted (13 of 18); the decision becomes linearly readable at argument emission but not at the examined pre-generation sites, and is not removed by the tested prefill- or decode-time steering, motivating argument-provenance controls.

↑