arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15200cs.CLcs.AIcs.LG

用于扩散语言模型的掩码感知策略梯度

Mask-Aware Policy Gradients for Diffusion Language Models

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Haran Raajesh, Kulin Shah, Adam Klivans, Philipp Krähenbühl

AI总结:

研究针对强化学习扩展到MDLMs的难题,将MDLM生成形式化为两阶段动作MDP,使策略梯度分解为令牌项和掩码项,结合优化这两项在数学推理和编码基准测试中取得了最优成绩。

AI中文摘要:

强化学习已被证明对改善大语言模型中的推理有效,但由于对数似然估计的难处理性,将其扩展到掩码扩散语言模型(MDLM)仍然具有挑战性。现有方法仅通过对令牌预测进行建模来近似此对数似然,忽略了生成过程中位置被解掩码的顺序。我们观察到MDLM生成在每个步骤涉及两个决策:在每个掩码位置放置什么令牌以及重新掩码哪些位置。我们将此形式化为两阶段动作MDP,表明策略梯度自然地分解为令牌项和掩码项。结合对这两个项的优化在数学推理和编码基准测试中产生了最先进的结果,在GSM8K上得分为87.1%,在MBPP上得分为53.4%。

英文摘要:

Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling only the token predictions, ignoring the order in which positions are unmasked during generation. We observe that MDLM generation involves two decisions at each step: what tokens to place at each masked position and which positions to remask. We formalize this as a two-stage action MDP, showing that the policy gradient naturally decomposes into a token term and a masking term. Combining optimization of both terms leads to state-of-the-art outcomes on mathematical reasoning and coding benchmarks, with scores of 87.1% on GSM8K and 53.4% on MBPP.

补充信息

↑