发表机构
HKUST(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对少步扩散语言模型现有离线策略蒸馏的轨迹不匹配问题,提出 OPTD 策略内转换蒸馏方法,在四个推理基准上改善质量效率权衡并取得最优质量约束 AUP。
AI 中文摘要
扩散语言模型(dLLM)可并行预测多个 token,但准确生成仍需大量迭代去噪步骤。少步蒸馏通过将多个教师步骤压缩为单个学生转换来加速解码。然而,现有方法在离线策略轨迹上构建监督。推理时,学生的早期并行承诺会改变后续预测的上下文,使其实际访问的状态偏离监督状态——而这正是步骤压缩最激进的时候。策略内蒸馏是解决这种不匹配的自然方法,但它留下了每个转换应推进多远的问题:仅匹配教师的下一个动作会限制压缩,而不加区分地合并未来动作会违反中间依赖关系。为解决这一限制,我们提出 OPTD(On-Policy Transition Distillation with Consistency-Guided Adaptive Compression,一致性引导自适应压缩的策略内转换蒸馏)。它从少步学生自身轨迹中采样部分状态,使用冻结的仅问题教师识别结果对齐的未来候选,并按当前状态置信度对其排序。然后选择最长前缀,其联合承诺保留教师的 rollout 结果。集合瓶颈目标将每个已验证的未来候选提升到解码器的释放阈值,而冻结教师的 KL 锚点则正则化所有其他活动位置。目标构建和训练均不使用黄金响应。在四个数学推理和代码生成基准上,OPTD 始终改善质量-效率权衡,并在评估的少步基线中获得最强的质量约束 AUP。
英文摘要
Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates decoding by compressing multiple teacher steps into a single student transition. However, existing methods construct supervision on off-policy trajectories. At inference, the student's early parallel commitments alter the context of later predictions, so the states it actually visits drift away from the supervised ones--precisely when step compression is most aggressive. On-policy distillation is a natural remedy for this mismatch, but it leaves open how far each transition should advance: matching only the teacher's next action limits compression, while indiscriminately merging future actions can violate intermediate dependencies. To address this limitation, we propose OPTD, On-Policy Transition Distillation with consistency-guided adaptive compression. It samples partial states from the few-step student's own trajectories, uses a frozen, question-only teacher to identify outcome-aligned future candidates, and orders them by current-state confidence. The method then selects the longest prefix whose joint commitment preserves the teacher's rollout outcome. A set-bottleneck objective promotes every verified future candidate to the decoder's release threshold, while a frozen-teacher KL anchor regularizes all other active positions. Neither target construction nor training uses a gold response. Across four mathematical reasoning and code-generation benchmarks, OPTD consistently improves the quality--efficiency trade-off and attains the strongest overall quality-constrained AUP among the evaluated few-step baselines.
Comments9 pages, 4 figures, 5 tables