保存你的饱和数据:在基于组的强化学习中超越奖励饱和进行学习
Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL
浏览论文内容
中文总结 AI 辅助
针对组相对强化学习中数据奖励饱和导致学习信号消失的问题,提出在展开生成层面引入高质量错误解作为负样本等干预策略,有效回收饱和数据,使GRPO性能提升6.4%-9.0%。
中文摘要 AI 辅助
组相对强化学习(Group-relative reinforcement learning, RL)依赖于采样响应之间的奖励差异来估计信息丰富的相对优势。随着语言模型能力日益增强,现有训练数据可能变得奖励饱和:同一问题的所有采样响应可能获得同等高的奖励,此时组相对学习信号消失,使得先前有用的数据变得过时。在本工作中,我们研究是否可以从这类饱和数据中恢复有用的学习信号。我们研究了组策略RL流水线的四个层面的干预措施——数据、展开(rollout)、奖励和优势——并仅在饱和推理数据上进行广泛的RL训练。虽然标准GRPO在饱和数据上几乎总是产生接近零的优势和接近噪声的信号,但多样化的干预措施成功回收并重新利用此类数据:在所提出的策略中,展开生成层面的干预措施始终最为有效:促使策略生成“高质量”的错误解决方案,将具有较差奖励的展开作为负样本引入饱和组,这使GRPO在Qwen3-1.7B和4B上的性能提升了6.4%至9.0%。其他干预措施如增加展开温度或添加辅助奖励也能恢复非零优势,但产生的收益不太一致。进一步分析表明,有效的负展开需要信息丰富的负轨迹,该方法在与非饱和数据共存时仍然有效,并支持对新饱和示例进行迭代回收。虽然越来越强的LLM会使更多数据变得饱和,但我们的结果表明,不要浪费你的饱和数据:通过正确的策略,它们可以在日益数据稀缺的世界中被回收为有用的RL训练信号。
英文摘要
Group-relative reinforcement learning (RL) relies on reward variation among sampled responses to estimate informative relative advantages. As language models become increasingly capable, existing training data can become reward-saturated: all sampled responses to the same problem might receive equally high rewards, where the group-relative learning signals vanish and leave previously useful data obsolete. In this work, we investigate whether useful learning signals can be recovered from such saturated data. We study interventions at four levels of group-policy RL pipelines---data, rollout, reward, and advantage---and conduct extensive RL training on saturated reasoning data only. While standard GRPO on saturated data would almost always yield near-0 advantages and near-noise signals, diverse interventions successfully recycle and repurpose such data: among the proposed strategies, interventions at rollout generation are consistently most effective: nudging the policy to generate ``high-quality'', incorrect solutions introduces rollouts with poor rewards into saturated groups as negative samples, which turns out to improve GRPO by 6.4% to 9.0% across Qwen3-1.7B and 4B. Other interventions such as increasing rollout temperature or adding auxiliary rewards can also restore non-zero advantages, but yield less consistent gains. Further analyses show that effective negative rollouts require informative negative trajectories, that the method remains effective alongside unsaturated data, and that it supports iterative recycling of newly saturated examples. While increasingly stronger LLMs would render more data as saturated, our results demonstrate that don't waste your saturated data: with the right strategies they can be recycled into useful RL training signals in an increasingly data-scarce world.