AI 中文总结
研究在标准结果奖励+GRPO设置下,无基于长度的加权方案能同时兼具梯度无偏性与长度不变性。通过参数族刻画权衡谱,揭示GRPO和Dr. GRPO各有优劣,处于不可避免的权衡两端,都非完美。
AI 中文摘要
组相对策略优化(GRPO)是用于训练大语言模型推理能力的主要强化学习算法,被DeepSeek - R1采用。近期改进的Dr. GRPO(COLM 2025)识别出GRPO中按轨迹长度归一化导致的响应级长度偏差,并提议去除该归一化,称所得优化器“无偏”。本文表明此说法不完整。具体建立了一个不可能性定理:在标准结果奖励+GRPO设置下,没有基于长度的加权方案能同时实现梯度无偏性(梯度估计器是真实策略梯度的无偏估计)和长度不变性(每个轨迹对梯度的有效贡献与其令牌长度无关)这两个属性。GRPO近似满足长度不变性但违反梯度无偏性;Dr. GRPO满足梯度无偏性但违反长度不变性。通过参数族$f_\alpha(L)=L^{\alpha - 1}$刻画了完整的权衡谱,其中$\alpha = 0$恢复GRPO,$\alpha = 1$恢复Dr. GRPO,并进行定量分析表明Dr. GRPO的长度偏差会使较长轨迹在梯度更新中占主导,其比例与长度比成正比。结果表明两种算法都并非完美,它们处于一个基本且不可避免的权衡的两端。
英文摘要
Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1. The recent improvement Dr. GRPO (COLM 2025) identifies the response-level length bias caused by per-trajectory length normalization in GRPO and proposes removing this normalization, claiming the resulting optimizer is "unbiased." We show that this claim is incomplete. Specifically, we establish an impossibility theorem: under the standard outcome reward + GRPO setting, no length-based weighting scheme can simultaneously achieve the following two properties. (P1) Gradient unbiasedness: the gradient estimator is an unbiased estimate of the true policy gradient. (P2) Length invariance: each trajectory's effective contribution to the gradient is independent of its token length. GRPO approximately satisfies P2 but violates P1; Dr. GRPO satisfies P1 but violates P2. We characterize the complete tradeoff spectrum via the parametric family f_alpha(L) = L^{alpha - 1}, where alpha = 0 recovers GRPO, alpha = 1 recovers Dr. GRPO, and provide quantitative analysis showing that Dr. GRPO's length bias can cause longer trajectories to dominate gradient updates by a factor proportional to the length ratio. Our results reveal that neither algorithm is universally "done right"; they occupy opposite ends of a fundamental and unavoidable tradeoff.