arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34509cs.LGcs.CL

低置信度重掩蔽限制灵活性:在扩散大语言模型中实现任意顺序生成的多样化展开潜力

Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs

Moongyu Jeon, Dongjae Jeon, Bumjun Kim, Mingyu Kim, Albert No

首次发表
浏览论文内容

中文总结 AI 辅助

本文发现掩蔽扩散语言模型多样性损失主要源于低置信度重掩蔽(LCR)的过滤机制,提出用最高概率位置选择(TPP)替代LCR,并引入熵引导初始化(EGI),从而恢复并提升生成多样性与覆盖度。

中文摘要 AI 辅助

掩蔽扩散语言模型支持任意顺序生成,这为产生多样化输出提供了一种自然途径。然而,近期研究认为这种灵活性反而通过延迟高不确定性标记(这些标记可能导致不同的生成路径)降低了多样性。我们将这种多样性损失追溯至并非任意顺序生成本身,而主要归因于低置信度重掩蔽(LCR),一种广泛使用的解码规则。在每一步中,LCR在每个掩蔽位置采样一个标记,但只提交具有最高概率的采样标记,过滤掉其余标记。我们证明,随着更多位置竞争,该机制可能指数级地抑制较低概率的标记,并在LLaDA中观察到同样的抑制现象。相比之下,最高概率位置选择(TPP)——其常与LCR在共享标签“基于置信度的解码”下被混淆——避免了这种多样性损失。TPP首先选择其最可能标记具有最高概率的位置,然后直接在该位置的分布中进行采样。用TPP替换LCR可恢复多样性,并产生与从左到右解码相当的Pass@$k$,这表明所报告的多样性损失主要源于LCR的过滤,而非先生成高置信度位置。为进一步利用顺序灵活性,我们引入了熵引导初始化(EGI),该方法在最高熵位置采样第一个标记,然后遵循TPP。这一简单修改进一步改善了展开多样性和解决方案覆盖率,超越了从左到右解码,其增益延伸到下游策略优化,突显了任意顺序生成在多样化展开中的潜力。

英文摘要

Masked diffusion language models support arbitrary-order generation, suggesting a natural way to produce diverse outputs. However, recent work argues that this flexibility reduces diversity by delaying high-uncertainty tokens that can lead to different generation paths. We trace this diversity loss not to arbitrary-order generation itself, but largely to low-confidence remasking (LCR), a widely used decoding rule. At each step, LCR samples a token at every masked position but commits only the sampled token with the highest probability, filtering out the rest. We show that this mechanism can exponentially suppress lower-probability tokens as more positions compete, and observe the same suppression in LLaDA. In contrast, top-probability position selection (TPP), which has often been conflated with LCR under the shared label confidence-based decoding, avoids this diversity loss. TPP first selects the position whose most likely token has the highest probability, then samples directly from that position's distribution. Replacing LCR with TPP restores diversity and yields Pass@$k$ comparable to left-to-right decoding, suggesting that the reported diversity loss stems largely from LCR's filtering rather than from generating high-confidence positions first. To further exploit order flexibility, we introduce Entropy-Guided Initialization (EGI), which samples the first token at the highest-entropy position and then follows TPP. This simple modification further improves rollout diversity and solution coverage beyond left-to-right decoding, with gains extending to downstream policy optimization, highlighting the potential of arbitrary-order generation for diverse rollouts.

发表机构

  • Yonsei University(延世大学)
  • KRAFTON AI(魁匠团AI)
  • Kookmin University(国民大学)

机构由 AI 辅助整理,请以论文原文为准。

↑