发表机构
Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory(上海交通大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对对比解码成本高的问题,提出解耦对比解码,采用专家对齐轻量级提案生成器与标准投机验证,实现显著加速并降低MMLU提案路径延迟。
AI 中文摘要
对比解码(CD)可提升生成质量,但其业余模型传递会使解码成本高昂。用投机解码加速CD会引发提案对齐问题:对比信号应塑造草稿生成器,还是仅保留在验证阶段?我们在轻量级特征级草稿生成器机制下研究该问题。两项受控诊断实验(匹配的Cross-alpha训练与近似双草稿生成器分解)得出相同结论:感知对比的草稿生成并不始终优于专家对齐的草稿生成,因为对比校正通常弱于草稿生成器误差,且重构会放大该误差。我们提出解耦对比解码(DCD),其采用专家对齐的轻量级提案生成器进行草稿生成,仅在不变的CD验证阶段使用业余模型。标准投机验证保留普通CD的输出分布。在主要8B设置下,基于EAGLE3的DCD相较于普通CD实现了1.65至1.95倍的平均贪心加速,且相对于业余耦合提案路径,将MMLU提案路径延迟降低了约5至12倍。
英文摘要
Contrastive Decoding (CD) improves generation quality, but its amateur-model pass makes decoding expensive. Accelerating CD with speculative decoding raises a proposal-alignment question: should the contrastive signal shape the drafter, or should it remain only in verification? We study this question in the lightweight feature-level drafter regime. Two controlled diagnostics, matched Cross-alpha training and an Approximate Dual-Drafter decomposition, give the same diagnosis: contrastive-aware drafting does not consistently improve over expert-aligned drafting because the contrastive correction is usually weaker than drafter error, and reconstruction can amplify that error. We introduce Decoupled Contrastive Decoding (DCD), which drafts with an expert-aligned lightweight proposer and applies the amateur only in unchanged CD verification. Standard speculative verification preserves the vanilla-CD output distribution. Across the main 8B settings, EAGLE3-based DCD achieves average greedy speedups of 1.65 to 1.95x over vanilla CD and reduces MMLU proposal-path latency by about 5 to 12x relative to amateur-coupled proposal paths.
Comments28 pages, 11 figures, 20 tables. Code: https://github.com/chadlzx/dcd Accepted to EMNLP 2026 (Main Conference)