arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

掩码扩散语言模型中的可靠并行解码

Reliable Parallel Decoding in Masked Diffusion Language Models

Zhenghao He, Bohan Liu, Guangzhi Xiong, Aidong Zhang

arXiv 2609.36452首次发表:更新:

发表机构

University of Virginia(弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对掩码扩散语言模型并行解码中预测不可靠的问题,提出无需训练的可靠并行解码(RPD)方法,基于逐层稳定性和置信度选择候选,在熵预算下提交,兼顾高吞吐量与准确性。

AI 中文摘要

掩码扩散语言模型(MDLMs)能够通过并行预测多个掩码标记来高效生成文本,但来自同一次前向传播的预测在同时提交时并不一定可靠。我们研究了何时并行提交是可靠的。我们的诊断表明,置信度本身并不能决定可靠的提交顺序:序列末尾的高置信度预测可以在其支持计算建立之前就确定答案,并且当下游预测的上游上下文不确定性增加时,下游预测的可靠性会降低。同时,单次前向传播已经可以解析多个掩码标记,并且在最终层中保持稳定的预测更可能是正确的。基于这些发现,我们提出了可靠并行解码(RPD),这是一种无需训练的方法,它根据逐层预测稳定性和最终置信度选择候选,并在其前面掩码位置的累积熵预算下提交它们。RPD 推迟具有不确定上游上下文的预测,同时并行提交其余候选,而不依赖固定的块调度。在 LLaDA 和 Dream 上的数学推理和代码生成基准测试中,RPD 在评估方法中实现了最高的解码吞吐量,同时保持或提高了准确性。

英文摘要

Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑