发表机构
Seoul National University; Interdisciplinary Program in Artificial Intelligence, Seoul National University; Artificial Intelligence Institute, Seoul National University; Department of Intelligence and Information, Seoul National University(首尔大学; 首尔大学人工智能交叉项目; 首尔大学人工智能研究院; 首尔大学智能与信息系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对扩散大语言模型强化学习中在线策略 rollout 稀缺的问题,提出 ERILS 方法整合外部策略 rollout,在数独等任务上显著提升性能,验证了 rollout 构建与奖励处理的重要性。
AI 中文摘要
近期针对扩散大语言模型(dLLMs)的强化学习方法,通常依赖目标 dLLM 自身生成的在线策略(on-policy)rollout。然而,当成功的在线策略 rollout 稀缺时,在线策略训练可能获得的正奖励极少,仅能取得有限进展。为缓解这一问题,我们探索将更强外部策略生成的高奖励 rollout,与目标 dLLM 的在线策略 rollout 相结合。但直接整合这些外部 rollout 会引入两个实际挑战:rollout 长度存在差异,以及联合处理在线策略与外部 rollout 的奖励时出现不稳定。为解决这些挑战,我们提出带长度控制与源特定处理的外部 rollout 整合方法(ERILS),该方法可控制外部 rollout 的长度,并分别处理在线策略与外部 rollout 的奖励。在数独(Sudoku)、倒计时(Countdown)和 MATH500 上进行的零样本评估实验显示,ERILS 在所有三个任务中均提升了多样本性能,其中在数独任务上的提升最大。在数独任务中,ERILS 实现了 98.4% 的 4 个样本中最优完成准确率,而最强基线仅为 40.3%;在 128、256 和 512 token 的生成长度下,ERILS 在数独任务上还保持了约 90% 的确定性单样本完成准确率。我们的组件分析进一步表明,长度受控的外部 rollout 比未受控的外部 rollout 更有效,且源特定奖励处理可避免联合奖励处理时出现的训练崩溃。这些结果表明,rollout 构建与奖励处理是将外部 rollout 整合入 dLLM 强化学习时的重要设计维度。
英文摘要
Recent reinforcement learning methods for diffusion large language models (dLLMs) commonly rely on on-policy rollouts generated by the target dLLM itself. When successful on-policy rollouts are scarce, however, on-policy training may receive little positive reward and make only limited progress. To mitigate this problem, we explore incorporating higher-reward rollouts generated by a stronger external policy alongside on-policy rollouts from the target dLLM. However, directly incorporating these external rollouts introduces two practical challenges: differences in rollout length and instability when jointly processing rewards from on-policy and external rollouts. To address these challenges, we propose External Rollout Integration with Length Control and Source-Specific Processing (ERILS), which controls external-rollout length and processes the rewards of on-policy and external rollouts separately. Experiments on Sudoku, Countdown, and MATH500 under zero-shot evaluation show that ERILS improves multi-sample performance across all three tasks, with the largest gains on Sudoku. On Sudoku, ERILS achieves 98.4% best-of-4 completion accuracy, compared with 40.3% for the strongest baseline. ERILS also maintains approximately 90% deterministic single-completion accuracy on Sudoku across generation lengths of 128, 256, and 512 tokens. Our component analysis further shows that length-controlled external rollouts are more effective than uncontrolled external rollouts, and that source-specific reward processing avoids the training collapse observed with joint reward processing. These results show that rollout construction and reward processing are important design dimensions when integrating external rollouts into dLLM reinforcement learning.