arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于微调离散扩散模型的连续时间强化学习框架

A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models

Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang

arXiv 2607.14522首次发表:更新:

发表机构

Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出连续时间强化学习框架,推导策略梯度方法得到PPO和GRPO的连续时间变体。以此开发框架微调离散扩散模型,无需奖励信号可微,能纳入中间奖励或优势信号,还为MDM提供统一视角,用轨迹子采样技术降低计算成本,在相关任务中验证了方法有效性。

AI 中文摘要

我们通过随机控制方法在具有离散状态空间和可能任意动作空间的连续时间内制定强化学习(RL),其中状态动态被建模为受控连续时间马尔可夫链(CTMC)。我们考虑策略优化问题并推导相应的策略梯度方法,得到近端策略优化(PPO)和群体相对策略优化(GRPO)的连续时间变体。作为主要应用,我们开发了一个完整的连续时间RL框架来微调基于分数的离散扩散模型。该框架实现奖励驱动的优化,无需奖励信号可微。与仅依赖终端奖励的现有基于GRPO的方法不同,我们的公式允许在去噪轨迹中纳入中间奖励或优势信号。专门针对掩码扩散模型(MDM)时,我们的框架包含词汇单纯形上丰富的策略参数化,概率比易于分析处理,为MDM中的探索和策略优化提供统一视角。对于掩码扩散大语言模型(dLLM),我们进一步提出轨迹子采样技术来有效估计计算成本高昂的轨迹似然,降低计算每个位置概率比的计算成本。我们在低维熵正则化优化问题以及dLLM在数学推理和编码任务上的RL后训练中展示了我们方法的有效性。

英文摘要

We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC). We consider policy optimization problems and derive corresponding policy gradient methods, leading to continuous-time variants of proximal policy optimization (PPO) and group relative policy optimization (GRPO). As a primary application, we develop a continuous-time RL framework for fine-tuning score-based discrete diffusion models, which enables reward-driven optimization without requiring differentiability on the reward signals. In contrast to the existing GRPO-based approaches that only rely on terminal rewards, our formulation allows intermediate reward or advantage signals to be incorporated throughout the denoising trajectory. Importantly, when specialized to masked diffusion models (MDMs), our framework encompasses a rich class of policy parameterizations over the vocabulary simplex with analytically tractable probability ratios, providing a unified perspective on exploration and policy optimization in MDMs. For diffusion large language models (dLLMs), we further propose trajectory subsampling techniques to efficiently estimate computationally prohibitive trajectory likelihoods, reducing the computational cost of computing per-position probability ratios. We showcase the effectiveness of our methods on both low-dimensional entropy-regularized optimization problems and RL post-training of dLLMs on reasoning and coding tasks.

Comments37 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑