发表机构
University of California-San Diego; Microsoft Research(加州大学圣迭戈分校; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从数学上分解富集轨迹修正的梯度误差为尺度、旋转和方差,提出序贯蒙特卡洛权重修正机制,并在OpenMathReasoning上验证其避免稀疏域崩溃的作用。
AI 中文摘要
当奖励稀疏时,使用可验证奖励的强化学习(RLVR)通常借助提示或中间指导来生成更多成功的轨迹。这种富集(enrichment)会通过重要性权重进行修正,否则会使策略梯度更新产生偏差;但现有方法为避免修正带来的高方差,要么省略修正,要么截断重要性权重。因此,在RLVR中,修正富集轨迹究竟带来何种得失,目前仍不清楚。我们从数学上证明,省略或截断修正会隐式地重新加权所定义的奖励,并将由此产生的梯度误差分解为尺度、旋转和方差三个部分,以解释它们对学习的各自影响。为使修正切实可行,我们提出一种新颖的序贯蒙特卡洛(SMC)权重修正机制,在稳定性和粒子顺序假设下,该机制能将标准修正方差随样本长度的指数级复合增长缓和为加性累积。随后,我们将富集分析应用于解释在OpenMathReasoning的稀疏子集上对Qwen3-1.7B进行微调的结果,从而确立一个具体机制,说明富集(无论是否修正)如何帮助避免稀疏领域中的崩溃。我们的主要贡献在于从总体上理解富集和修正如何影响训练,而非声称任何一种操作模式优于未富集的RL。
英文摘要
When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in order to avoid the high variance of correction. It thus remains unclear what exactly is gained or lost in RLVR by correcting enriched rollouts. We show, mathematically, that omitted or truncated correction implicitly reweights the defined reward, and we decompose the resulting gradient error into scale, rotation, and variance to explain their distinct effects on learning. To make correction practical, we develop a novel sequential Monte Carlo (SMC) weight correction mechanism that, under stability and particle-order assumptions, tempers the exponential compounding of standard correction variance over the length of a sample to an additive accumulation. We then apply our analysis of enrichment to interpreting the results of a fine-tune of Qwen3-1.7B on a sparse band of OpenMathReasoning, establishing a concrete mechanism of how enrichment, both corrected and uncorrected, helps avoid collapse in sparse domains. Our main contribution is to understand, in general, how enrichment and correction can affect training, as opposed to claiming that either mode of operation is superior to unenriched RL.