arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Fork-dLLM:避免扩散语言模型中的灵活性陷阱

Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models

Stipe Frković, Metod Jazbec, Christian A. Naesseth

arXiv 2609.39859首次发表:更新:

发表机构

University of Amsterdam; UvA-Bosch Delta Lab, University of Amsterdam(阿姆斯特丹大学; 阿姆斯特丹大学-博世三角洲实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散语言模型因高熵分叉位置导致生成多样性和后训练收益受限的问题,提出混合采样器Fork-dLLM及后训练方法ForkGRPO,在保持并行生成效率的同时,匹配自回归采样的性能并降低训练成本。

AI 中文摘要

掩码扩散语言模型(dLLMs)在与基于置信度的采样器结合时,通过并行令牌生成展现出实现更快推理的强大潜力。然而,近期研究表明,此类方法可能会推迟对存在多种合理续写的高熵分叉位置进行去掩码处理。这导致生成多样性降低,表现为pass@k缩放效果变差,并限制了从强化学习后训练中获得的收益。为避免这一灵活性陷阱,先前的工作提倡使用自回归(AR)采样。在此,我们表明放弃基于置信度的采样并无必要,且一旦考虑推理成本,这种做法是浪费的。我们首先提出Fork-dLLM,一种简单的混合采样器,仅在不确定的回退步骤中使用AR式排序,而在其他情况下保留并行生成。随后,我们将相同原则扩展到后训练,提出ForkGRPO,它使用Fork-dLLM的轨迹,并仅在回退步骤应用GRPO目标,在大幅降低轨迹生成和优化成本的同时,保持精确的策略似然比。在我们的实验中,Fork-dLLM在匹配AR采样强大的pass@k缩放效果的同时,效率提高了2-3倍;ForkGRPO在显著降低训练成本的情况下,实现了与基于AR的GRPO基线相当或更优的下游性能。

英文摘要

Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training. To avoid this flexibility trap, prior work advocated for autoregressive (AR) sampling. Here, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. We first propose Fork-dLLM, a simple hybrid sampler that uses AR-style ordering only at uncertain fallback steps while retaining parallel generation otherwise. We then extend the same principle to post-training with ForkGRPO, which uses Fork-dLLM rollouts and applies the GRPO objective only at fallback steps, preserving exact policy-likelihood ratios while substantially reducing rollout and optimization cost. In our experiments, Fork-dLLM matches the strong pass@k scaling of AR sampling while being 2-3x more efficient, and ForkGRPO achieves downstream performance comparable to or better than AR-based GRPO baselines at a substantially lower training cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑