CForce:通过一致性强制提升扩散大语言模型(dLLMs)的并行解码性能
CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing
- Shanghai Jiao Tong University(上海交通大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出CForce方法,通过一致性强制提升dLLMs的并行解码性能,在LLaDA模型上实验证实其在高并行解码预算下可优化速度-质量权衡。
AI中文摘要:
扩散大语言模型(dLLMs)通过单次前向传播预测多个掩码来加速语言生成,但现有dLLMs在激进的并行解码策略下,早期去噪阶段的预测可能不可靠,导致错误传播至后续阶段。为解决该问题,本文提出适用于dLLMs的一致性强制(Consistency Forcing,CForce),这是一种蒸馏方法,用于强制早期阶段的掩码预测与后期阶段对齐。CForce在预先收集的自回滚轨迹上训练模型,从而提升训练-推理对齐性。本文引入置信度自适应KL散度作为蒸馏目标,结合前向和反向KL的优势;还对一致性目标进行理论分析,解释CForce为何能近似最小化早期阶段的预测误差。关键的是,该公式既适用于掩码到令牌的解码,也适用于支持编辑的解码;在支持编辑的场景中,后期的令牌到令牌细化可为早期掩码状态预测提供额外监督。在非编辑和支持编辑的LLaDA模型上开展的实验表明,该方法在速度-质量权衡上有所提升,尤其在高并行解码预算下效果显著。代码可从指定URL获取。
英文摘要:
Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. However, existing dLLMs can suffer from unreliable predictions in early denoising stages under aggressive parallelism strategies, leading to errors that can propagate to later stages. To tackle this issue, we present Consistency Forcing (CForce) for dLLMs, a distillation method to force the mask predictions of early stages to align with those of later stages. CForce trains the model on pre-collected self-rollout trajectories, thereby improving training-inference alignment. We introduce Confidence Adaptive KL Divergence as a distillation objective to conjoin the merits of forward and reverse KL. We further provide a theoretical analysis for the consistency objective to explain why CForce can approximately minimize the prediction error of early stages. Critically, the same formulation applies to both mask-to-token decoding and edit-capable decoding; in the edit-capable case, later token-to-token refinements provide additional supervision for earlier masked-state predictions. Experiments on non-edit and edit-capable LLaDA models show improved speed-quality trade-offs, especially under high-parallelism decoding budgets. Code is available at: https://github.com/inclusionAI/dFactory.