arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推测解码中用于接受崩溃的对抗性提示

Adversarial Prompts for Acceptance Collapse in Speculative Decoding

Run Wang, Chaoyi Zhou, Xi Liu, Yi Zhu, Amir Salarpour, Pedram MohajerAnsari, Zhi-Qi Cheng, Feng Luo, Siyu Huang, Mert D. Pesé

arXiv 2607.21804首次发表:更新:

发表机构

Clemson University; Wayne State University; University of Washington(克莱姆森大学; 韦恩州立大学; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对推测解码中草稿-目标对齐漏洞的攻击,提出ADSD方法,利用Soft-Collapse及目标保留目标生成对抗性后缀,在GSM8K数据集上验证其有效性,还证明该漏洞存在于多方面。

AI 中文摘要

无损加速方案,如推测解码,通过依赖草稿模型和目标模型之间的动态令牌级对齐来显著加快推理速度。然而,这种语义等价性的保证掩盖了一个严重的操作漏洞:草稿-目标对齐可能会受到系统性攻击。在本文中,我们引入了ADSD,据我们所知,这是第一种通过将草稿概率质量推向目标模型不太可能接受的令牌来使验证器接受崩溃的提示后缀攻击。ADSD使用Soft-Collapse,一种从不对称推测接受规则派生的验证器对齐代理,以及一个防止明显任务损坏的目标保留目标。ADSD成功生成了高效的对抗性后缀。在GSM8K数据集上,我们的攻击在保持任务质量的同时将平均采样时间增加了62.3%。我们进一步表明,这种漏洞存在于不同领域、推测解码策略和模型架构中。

英文摘要

Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence masks a severe operational vulnerability: draft-target alignment can be systematically attacked. In this paper, we introduce ADSD, which, to the best of our knowledge, is the first prompt-suffix attack that collapses verifier acceptance by pushing draft probability mass toward tokens the target is unlikely to accept. ADSD uses Soft-Collapse, a verifier-aligned surrogate derived from the asymmetric speculative acceptance rule, together with a target-preservation objective that discourages obvious task corruption. ADSD successfully generates highly effective adversarial suffixes. On the GSM8K dataset, our attack increases the mean sample time by 62.3% while preserving the task quality. We further show that this vulnerability exists across different domains, speculative decoding strategies, and model architectures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑