AI 中文总结
该研究针对GRPO中正确解重复导致的稀有解信用不足问题,提出基于Strategy Cues的Cue-GRPO方法,在Qwen2.5-Math-7B等模型上提升AIME重复采样性能,仅增加少量训练开销。
AI 中文摘要
带可验证奖励的强化学习(RLVR)通常将每个正确完成过程作为独立学习信号进行优化。在GRPO中,这种完成级别的均匀性会产生结构级偏差:重复出现的正确解形式会根据其被采样的频率积累正系数质量,而稀有形式则获得有限的信用。我们将这种行为形式化为多诱导的结构级信用集中,并引入一种分区条件规则,该规则根据聚类稀有度重新分配正优势。Cue-GRPO通过使用确定性策略线索(Strategy Cues)构建已验证正确轨迹的rollout局部分区,无需辅助模型推理即可实例化此规则。在Qwen2.5-Math-7B和Llama-3.1-8B-Instruct上,Cue-GRPO提高了AIME重复采样性能,在高采样预算下增益最大。在评判分区(JP)下的信用再分配(CR)进一步表明,所提出的再分配机制可使用评判派生分区运行。Cue-GRPO仅比GRPO增加6%的挂钟训练开销。这些结果支持结构级信用再分配作为RLVR的实用设计轴,其中Strategy Cues为竞赛数学提供了低开销实现。代码可在https URL Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO获取。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GRPO, this completion-level uniformity creates structure-level skew: recurring correct solution forms accumulate positive coefficient mass in proportion to how often they are sampled, while rare forms receive limited credit. We formalize this behavior as multiplicity-induced structure-level credit concentration and introduce a partition- conditioned rule that redistributes positive advantages accord- ing to cluster rarity. Cue-GRPO instantiates this rule with- out auxiliary-model inference by using deterministic Strategy Cues to construct rollout-local partitions of verified-correct traces. Across Qwen2.5-Math-7B and Llama-3.1-8B-Instruct, Cue-GRPO improves AIME repeated-sampling performance, with the largest gains at high sampling budgets. Credit Re- distribution (CR) under Judge Partitions (JP) further indi- cates that the proposed redistribution mechanism can oper- ate with judge-derived partitions. Cue-GRPO adds only 6% wall-clock training overhead over GRPO. These results sup- port structure-level credit redistribution as a practical design axis for RLVR, with Strategy Cues providing a low-overhead implementation for competition mathematics. Code is avail- able at https://github.com/CzZ12/When-Correct-Solutions- Repeat-Rarity-Aware-Credit-Redistribution-for-GRPO.