发表机构
The Pennsylvania State University; Duke University(宾夕法尼亚州立大学; 杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究首次系统揭示投机解码的安全-效用不对称性,并提出SecureSD方法,通过对早期草稿令牌施加更严格验证,在保持效率与效用下显著提升安全性。
AI 中文摘要
投机解码通过首先使用一个较小的模型(称为草稿模型)生成候选令牌,然后由目标模型(即大型语言模型)进行验证以接受或拒绝,从而加速目标模型的推理。先前的研究主要关注投机解码的效率-效用权衡,例如有损投机解码,而其安全影响在很大程度上未被探索。在这项工作中,我们通过首次对投机解码的安全影响进行系统性研究来弥补这一空白。通过大规模测量研究,我们揭示了显著的安全-效用不对称性:在广泛的有损投机解码方法中,推理效率的提升以不成比例的高安全成本为代价,越狱和提示注入攻击的成功率增长速度远快于效用的下降。随后,我们提出了SecureSD,一种新的理论指导的投机解码方法,在保持效率和效用的同时增强安全性。具体而言,我们的理论分析表明,安全退化主要源于草稿模型生成的早期令牌。受此见解的启发,SecureSD在解码的早期位置对草稿模型令牌应用更严格的验证标准。在安全性和效用基准上的大量实验表明,与现有投机解码方法相比,SecureSD在保持效率和效用的同时显著提高了安全性。
英文摘要
Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emph{target model}, by first using a smaller model, referred to as the \emph{draft model}, to generate candidate tokens and then verifying them with the target model for acceptance or rejection. Prior studies primarily focused on the efficiency-utility trade-off of speculative decoding, e.g., lossy speculative decoding, leaving its security implications largely unexplored. In this work, we bridge this gap by providing the \emph{first} systematic study of the security implications of speculative decoding. Through a large-scale measurement study, we reveal a pronounced security-utility asymmetry: across a wide range of lossy speculative decoding methods, improvements in inference efficiency come at a disproportionately high cost to security, with attack success rates for jailbreak and prompt injection attacks increasing much faster than utility degrades. We then propose SecureSD, a new theory-guided speculative decoding method that enhances security while maintaining efficiency and utility. Specifically, our theoretical analysis reveals that security degradation primarily originates from the early tokens generated by the draft model. Motivated by this insight, SecureSD applies a stricter verification criterion to draft-model tokens at early decoding positions. Extensive experiments on both security and utility benchmarks demonstrate that SecureSD significantly improves security while preserving efficiency and utility compared to existing speculative decoding methods.
Comments18 pages, accepted by IEEE S&P 2027