推测校正:扩散语言模型的草稿后精炼解码
Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models
- National University of Singapore(新加坡国立大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对扩散语言模型提出草稿后精炼解码方案,通过 Flash-Flash 和 Mini-Flash 配置在 GSM8K、MBPP、MATH 数据集上提升准确率并加快推理速度,为双向精炼及快速生成提供新路径。
AI中文摘要:
扩散语言模型(DLMs)可双向修正 token,但标准解码流程常通过逐块生成文本将其适配为从左到右的生成模式。本文研究一种简单的即插即用推理模式:先生成完整草稿,再利用双向扩散对完整响应进行精炼。我们使用 LLaDA2.1-Flash 和 LLaDA2.1-Mini 评估两种配置:Flash-Flash 中,相同的 Flash 模型同时作为草稿生成器和精炼器,用于测试现有模型能否通过全局精炼改进自身的块自回归输出;Mini-Flash 中,受推测解码启发,我们提出推测校正:Mini 生成完整响应草稿,Flash 将其作为可编辑初始化进行修正。Flash-Flash 将 GSM8K-384 的准确率从 0.848 提升至 0.899,同时比所选的 Flash 块自回归基线快 1.20 倍,将 MBPP-384 的准确率从 0.545 提升至 0.693。延迟窗口匹配的仅 Flash 对照组表明,这些增益在对块自回归解码进行针对性调整后依然存在。因果消融实验显示,完整草稿可提供有用的初始化:从完全掩码的跨度进行精炼表现不佳,完整全局精炼在 GSM8K 上提供了明显的额外增益,而局部精炼在 MBPP 和 MATH 上捕获了大部分增益。Mini-Flash 提供了有用的质量-延迟权衡,包括 MATH-384 的性能为 0.294,而 Flash 的性能为 0.300,同时运行速度快 2.17 倍。这些结果支持帕累托前沿的解释,而非异构级联均匀匹配 Flash 质量的说法。总体而言,同模型的草稿-精炼方案证明双向精炼是 DLMs 有用的解码原语,而推测校正展示了一种无训练的快速 DLM 生成路径。
英文摘要:
Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion. Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations. In Flash-Flash, the same Flash model serves as both drafter and refiner, testing whether an existing model can improve its own block-autoregressive output through global refinement. In Mini-Flash, inspired by speculative decoding, we introduce speculative correction: Mini drafts a full response, and Flash revises it as an editable initialization. Flash-Flash improves GSM8K-384 accuracy from 0.848 to 0.899 while running 1.20 times faster than the selected Flash block-autoregressive baseline, and improves MBPP-384 from 0.545 to 0.693. Latency-window-matched Flash-only controls indicate that these gains persist after targeted tuning of block-autoregressive decoding. Causal ablations indicate that completed drafts provide useful initializations: refinement from a fully masked span performs poorly, full global refinement provides a clear additional gain on GSM8K, and local refinement captures much of the gain on MBPP and MATH. Mini-Flash provides useful quality-latency trade-offs, including MATH-384 performance of 0.294 versus 0.300 for Flash while running 2.17 times faster. These results support a Pareto-frontier interpretation rather than the claim that the heterogeneous cascade uniformly matches Flash quality. Overall, same-model draft-and-refine provides evidence that bidirectional refinement is a useful decoding primitive for DLMs, while speculative correction demonstrates a training-free route to fast DLM generation.