BASIS:基于预填充注意力探测的漏洞感知选择性提示注入防护
BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes
浏览论文内容
中文总结 AI 辅助
BASIS是一种基于预填充注意力探测的提示注入防御方法,通过级联门控的稀疏线性探测器实现精准检测,可在保持高注入检测率的同时减少不必要的过度拒绝。
中文摘要 AI 辅助
提示注入是大语言模型(LLM)应用中的关键安全威胁,攻击者通过在用户或外部数据中嵌入恶意指令劫持模型行为。现有检测方法仅能检测到注入的存在并在检测到后拒绝响应,却忽略了一个事实:对于许多现代对齐模型,精心设计的指令可抵御大多数注入攻击,这意味着不同指令和模型的注入鲁棒性存在显著差异,进而导致广泛存在不必要的过度拒绝:包含模型本可正确处理的注入的输入被错误拒绝。为解决这一过度拒绝问题,我们提出BASIS(鲁棒性感知提示注入防御),该防御方法使用注意力竞争比率(ρ)作为特征来训练两个稀疏线性探测器:存在探测器和漏洞探测器,两个探测器通过级联门控做出防御决策,无需额外的LLM推理。BASIS包含三个阶段:注入存在检测、单样本漏洞预测和指令鲁棒性评估;在线级联仅在模型实际会被破坏时才拒绝,从而避免对鲁棒指令的过度拒绝。在四个任务和六个开源LLM上的实验表明,BASIS在保持近乎完美的注入检测能力的同时,大幅降低了安全攻击样本的过度拒绝,尤其是在鲁棒指令模板下。
英文摘要
Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data. Existing detection methods only detect the presence of injection and refuse to respond upon detection, overlooking the fact that for many modern aligned models, well-crafted instructions can resist most injection attacks. This means that the injection robustness varies significantly across instructions and models. This leads to widespread unnecessary over-refusal: inputs containing injections that the model could have handled correctly are rejected incorrectly. To deal with this over-refusal issue, we propose BASIS (Robustness-Aware Prompt Injection Defense). This defense method uses the Attention Competition Ratio ($ρ$) as features to train two sparse linear probes: an existence probe and a breach probe. Both probes make defense decisions through cascaded gating, which does not require additional LLM inference. BASIS comprises three stages: injection existence detection, per-sample breach prediction, and instruction robustness assessment; the online cascade refuses only when the model would actually be compromised and thus avoids over-refusal on robust instructions. Experiments across four tasks and six open-source LLMs show that BASIS maintains near-perfect injection detection while substantially reducing over-refusal on safe attack samples, especially under robust instruction templates.
发表机构
- City University of Macau(澳门城市大学)
- Qilu University of Technology(齐鲁工业大学)
机构由 AI 辅助整理,请以论文原文为准。