发表机构
Imperial College London(帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出针对推测解码的新型推测拒绝攻击,通过附加对抗性后缀减慢LLM推理速度、增加成本,两种攻击方法可跨模型迁移,揭示了推测解码的攻击面。
AI 中文摘要
推测解码是一种流行的技术,通过在一次目标模型前向传播中验证多个草稿token,来提高大语言模型(LLM)推理的速度并降低成本。该技术的效果取决于草稿模型近似目标模型分布的能力。在本研究中,我们提出了推测拒绝攻击(Speculative Rejection Attacks, SRAs),这是一类新型攻击,可导致草稿模型与目标模型的一致性降低,使得每个草稿周期中被接受的草稿token减少。这会导致生成每个token所需的目标模型前向传播次数增加,从而减慢推理速度并增加受害者的成本。我们提出了两种攻击方法:在攻击者控制的内容后附加对抗性后缀,以降低受害者提示下的推测解码效果。两种攻击均优化了被接受的推测前缀的预期长度,分别通过目标模型对草稿提议的概率估计各深度的接受情况(Speedbump-P),或通过草稿与目标分布的重叠进行估计(Speedbump-D)。在某些情况下,这些攻击会使推测解码的速度降至比自回归解码更慢的程度。这种性能下降会降低输出质量;正则化可恢复输出质量,但会放弃大部分攻击效果,以隐蔽性为代价换取有效性的降低。此外,这些后缀在采样下仍保持有效,且可跨草稿模型(Speedbump-P)或跨共享同一草稿模型的目标模型(Speedbump-D)迁移。这些发现表明,推测解码的草稿-目标交互是一个现实的攻击面,攻击者可通过该攻击面生成对抗性输入来增加推理成本。
英文摘要
Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass. The resulting benefit depends on the ability of the drafter to approximate the target model's distribution. In this work, we study Speculative Rejection Attacks (SRAs), a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle. This leads to more target model forward passes needed per generated token, slowing down inference and increasing costs for the victim. We introduce two attacks which append an adversarial suffix to attacker-controlled content to degrade speculative decoding on a victim's prompts. Both attacks optimise the expected length of the accepted speculative prefix, estimating per-depth acceptance from the target's probability of the drafted proposals (Speedbump-P) or from the overlap between the draft and target distributions (Speedbump-D). In some cases, attacks degrade speculative decoding to the point of being slower than autoregressive decoding. The degradation reduces the output quality - regularisation restores output quality but gives up most of the degradation, trading effectiveness for stealthiness. Additionally, the suffixes remain effective under sampling, and transfer across drafters (Speedbump-P) or across target models sharing a drafter (Speedbump-D). These findings identify the draft-target interaction of speculative decoding as a realistic attack surface through which adversarial inputs can inflate inference costs.