发表机构
Indiana University; University of California, Riverside(印第安纳大学; 加州大学河滨分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出StepRS-GRPO,通过逐步调整风险系数改进扩散语言模型的强化学习,提升数学推理的准确率、覆盖率和答案多样性。
AI 中文摘要
扩散大语言模型(dLLMs)通过去噪序列或连续块来生成文本,从而允许并行揭示多个令牌。带可验证奖励的强化学习(RLVR)在这些决策中重用终端反馈,即使其条件上下文发生变化。我们提出了逐步风险敏感GRPO(StepRS-GRPO),它在去噪状态之间变化组优势变换的风险系数,同时保留底层训练器。对于二元奖励,我们证明该变换恰好是中心化结果优势的提示和状态相关重缩放。基于能力的校准建议了系数尺度,而端点和插值消融指导调度选择。在多个dLLM骨干和数学推理基准上,StepRS-GRPO在pass@1准确率和pass@k覆盖率上均优于中心化GRPO,同时增加了答案多样性。在我们的消融研究中,质量匹配对照支持状态分配和调度方向的贡献,并且在将优势的均方根(RMS)与中心化GRPO匹配后,增益仍然存在。推理轨迹诊断进一步表明,StepRS-GRPO带来的多样性增益超越了最终答案字符串。
英文摘要
Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states while retaining the underlying trainer. For binary rewards, we show that this transformation is exactly a prompt- and state-dependent rescaling of centered outcome advantages. A capability-based calibration suggests a coefficient scale, while endpoint and interpolation ablations guide schedule selection. Across multiple dLLM backbones and mathematical reasoning benchmarks, StepRS-GRPO improves both pass@1 accuracy and pass@k coverage over centered GRPO, while increasing answer diversity. In our ablation studies, mass-matched controls support the contributions of state allocation and schedule direction, and the gains persist after matching the root mean square (RMS) of the advantages to that of centered GRPO. Reasoning-trace diagnostics further show that the diversity gains from StepRS-GRPO extend beyond final-answer strings.
Comments40 pages, 12 figures, 2 tables. The first two authors contributed equally