先给出答案,后进行推理:扩散大语言模型中的承诺顺序
Answer First, Reason Later: When Commitment Order Costs Accuracy in Diffusion Language Models
- Graduate School of Data Science, Seoul National University(首尔国立大学数据科学研究生院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究发现扩散大语言模型(dLLMs)的任意顺序提交机制会导致推理失败,通过前沿门控承诺干预可恢复推理性能,同时保留并行解码优势。
AI中文摘要:
掩码扩散语言模型(dLLMs)可以按任意顺序提交 token,这一自由被宣传为其相对于自回归解码的核心优势。我们表明,在推理任务上,这种自由反而是失败的关键因素。在 GSM8K 数据集上对 LLaDA-8B 解码过程中的每一次提交进行记录,我们发现无约束(纯)解码会在轨迹的 15-24% 处提交最终答案,而此时仍有一半的推理区域处于掩码状态;随着画布增大,多达 90% 的问题会崩溃为仅输出答案。原因并非模型的终止信念——EOS(结束符)的“压力”在各类解码器中几乎相同——而是可达性:采样器是否能在较远位置作用于这些信念。一项 2×2 提示-解码器设计显示,思维链仅在有序承诺下才有帮助(交互提升 34.8 个百分点,95% 置信区间 [26.8, 42.8];无推理文本时解码器无差异),我们将这种交互分解为崩溃通道和顺序通道,并在 Dream-7B 和 MATH-500 上复现。一项单旋钮干预——前沿门控承诺——可因果性地恢复全部差距(从 0.528 到 0.852),同时保留高达 4 倍的并行解码,其测量的最优窗口在完全细化时从 w=1 翻转到每步 8 token 时无约束。我们的结果重新定义了现有的窗口式采样器,此前其动机为效率,如今则是解决其从未设计应对的推理缺陷的最小方案。
英文摘要:
Masked diffusion language models revise many masked output positions in parallel. We call a token committed once it becomes visible and is never masked again, and call a response answer-first when the final answer commits before the reasoning printed ahead of it. On 1,069 GSM8K test questions, an explicit step-by-step instruction increases the accuracy difference between unrestricted decoding and a decoder that permits commitment only near the left-most unresolved position; unrestricted decoding also produces more answer-first trajectories. On MATH-500, the two LLaDA models spend most of a short output canvas on reasoning that commits after the answer, and the benefit of frontier gating decreases as that postanswer writing disappears. Dream-7B has little post-answer writing and follows a different accuracy pattern. A controlled four-option task reserves a one-token answer position before generation. Delaying that position outperforms an equally timed reasoning-token delay on LLaDA-8B, LLaDA-1.5, and Dream-7B. The raw difference is largest on Dream, whose free accuracy on the controlled task is lower. Answers commit much earlier under the reserved-position interface than in ordinary free-form generation, which limits how far the intervention result can be generalized. Commitment order affects the context used to complete a response and the allocation of a finite output canvas.