SpecGuard:推理时后门检测,零额外成本
SpecGuard: Inference-Time Backdoor Detection For Free
浏览论文内容
中文总结 AI 辅助
SpecGuard利用投机解码的验证过程,在零额外计算成本下检测LLM推理时的后门触发行为,可靠识别隐蔽攻击,避免现有方法的额外生成开销。
中文摘要 AI 辅助
大型语言模型通常经过微调、共享或从第三方下载,因此部署的模型可能携带隐藏的后门,该后门在良性输入上表现正常,但当出现秘密触发器时,会切换到攻击者控制的行为。虽然后门可以在部署前进行审计,但对于频繁更新的模型,运行时监控仍然很重要。挑战在于LLM服务对延迟敏感:现有的推理时检测器要么依赖于对触发器形式的假设,这可能在隐蔽攻击上失效,要么需要额外的模型计算,例如输入扰动或额外的生成过程。我们引入了SpecGuard,一种推理时后门检测器,它以零额外模型计算成本重新利用投机解码。投机解码通过使用小型草稿模型提出令牌并由目标模型验证来加速推理。我们观察到,这一验证过程已经暴露了一个有用的信号:当后门被触发时,目标模型转向攻击者的行为,而干净的草稿模型不会预测这种转变,导致草稿令牌接受率发生变化。我们形式化了该信号出现的条件,并表明试图抑制该信号的攻击者必须同时削弱后门。在各种后门类型和模型家族中,SpecGuard可靠地检测触发行为,包括输入级过滤器无法察觉的隐蔽情况,同时避免了现有运行时检测器的额外生成成本。因此,投机解码可兼作检测后门LLM行为的免费、始终开启的信号。
英文摘要
Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger appears. While backdoors can be audited before deployment, runtime monitoring remains important for models that are frequently updated. The challenge is that LLM serving is latency-sensitive: existing inference-time detectors either rely on assumptions about the trigger form, which can fail on stealthy attacks, or require extra model computation, such as input perturbations or an additional generation pass. We introduce SpecGuard, an inference-time backdoor detector that repurposes speculative decoding at zero added model-computation cost. Speculative decoding speeds up inference by using a small draft model to propose tokens and a target model to verify them. We observe that this verification process already exposes a useful signal: when a backdoor is triggered, the target model shifts toward the attacker's behavior, while a clean draft model does not predict this shift, causing the draft-token acceptance rate to change. We formalize when this signal appears and show that an attacker who suppresses it must also weaken the backdoor. Across diverse backdoor types and model families, SpecGuard reliably detects triggered behavior, including stealthy cases where input-level filters are blind, while avoiding the extra generation cost of existing runtime detectors. Speculative decoding therefore doubles as a free, always-on signal for detecting backdoored LLM behavior.
发表机构
- Institute of Science Tokyo(东京科学大学)
- Microsoft Security Response Center(微软安全响应中心)
- Microsoft Azure(微软Azure)
- Shandong University(山东大学)
机构由 AI 辅助整理,请以论文原文为准。