发表机构
Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CanvasAnneal通过课程式教师引导注入与逐步退火,缓解扩散语言模型强化学习的探索瓶颈,在数学推理和工具使用基准上优于标准方法并加速收敛。
AI 中文摘要
扩散语言模型(DLMs)提供了有前景的并行生成能力,但在复杂推理和工具使用任务上落后于自回归模型。虽然强化学习(RL)最近已被应用于增强扩散语言模型,但标准的RL方法存在探索瓶颈。为了解决这个问题,我们从更强的教师模型中注入推理先验来引导RL探索。在本文中,我们介绍了CanvasAnneal,一个课程引导的扩散RL框架。在初始RL阶段,我们通过将教师生成的推理轨迹注入初始扩散画布来预热探索。随着训练的进行,我们逐渐移除这种引导,并要求模型更独立地生成推理轨迹。在数学推理和工具使用基准上,CanvasAnneal在MATH500、Countdown和Tau2上优于标准的diffu-GRPO,并在多个任务上显著加速了奖励提升,尽管收益因任务而异。我们的结果表明,结构化的训练时引导可以缓解扩散RL中的探索瓶颈,并加速在更困难任务上的收敛。
英文摘要
Diffusion Language Models (DLMs) offer promising parallel generation capabilities but lag behind autoregressive models in complex reasoning and tool-use tasks. While Reinforcement Learning (RL) has recently been applied to enhance DLMs, standard RL approaches suffer from an exploration bottleneck. To address this, we inject reasoning priors from a stronger teacher model to guide RL exploration. In this paper, we introduce CanvasAnneal, a curriculum-guided diffusion RL framework. During the initial RL phase, we warm-start exploration by injecting teacher-generated reasoning traces into the initial diffusion canvas. As training progresses, we gradually remove this guidance and require the model to generate more of the reasoning trajectory independently. Across mathematical reasoning and tool-use benchmarks, CanvasAnneal improves over standard diffu-GRPO on MATH500, Countdown, and Tau2 and substantially accelerates reward improvement on several tasks, while gains are task-dependent. Our results suggest that structured training-time guidance can alleviate exploration bottlenecks in diffusion RL and speed up convergence on harder tasks.
Comments16 pages, 2 figures, 6 tables