发表机构
State Key Laboratory for Novel Software Technology, Nanjing University; Institute of Wireless Communications Technology, Shanghai Jiao Tong University(南京大学现代软件技术国家重点实验室; 上海交通大学无线通信技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长思维链推理中均匀分配 token 权重的缺陷,提出反事实敏感性权重重分配方法,在数学推理基准上优于 GRPO,验证了特权诱导方向不可靠的诊断。
AI 中文摘要
带可验证奖励的强化学习(RLVR)是提升大语言模型长思维链(Long-CoT)推理能力的核心。GRPO 等无批评者方法将响应级奖励转换为优势值,并均匀分配给所有 token,忽略了它们对最终结果的贡献存在差异。在线自蒸馏(OPSD)则通过最小化非特权策略与特权自教师之间的前向 KL 散度,提供密集分布监督,隐含假设由此产生的似然变化编码了可靠的答案对齐信息。我们通过固定每个采样轨迹并在两种对立结果条件下重新评分来检验这一前提:一种断言正确,另一种断言错误。多数受影响 token 在两种条件下向同一方向变化,符号反转极少,且诱导的优化信号存在大量重叠;大的变化还集中在高度可替换的表层形式 token 上,而承载问题特定推理内容的 token 敏感性较低。这些发现表明,特权变化无法提供可靠的答案对齐方向,其幅度主要反映反事实敏感性而非 token 级学习价值。基于这些观察,我们提出反事实敏感性权重重分配(CSCR),这是 GRPO 的简单扩展,它降低高敏感 token 的权重并重新归一化 token 级优势值,以保留原始权重预算和验证器确定的方向。在长思维链数学推理基准上,CSCR 在相同策略更新次数下始终优于 GRPO 基线;针对性消融进一步证实了我们的诊断:特权诱导的方向不可靠,适度降权最有效,更强的调制会破坏优化稳定性。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their unequal contributions to the final outcome. On-policy self-distillation (OPSD) instead provides dense distributional supervision by minimizing the forward KL divergence between an unprivileged policy and a privileged self-teacher, implicitly assuming that the resulting likelihood shifts encode reliable answer-aligned information. We test this premise by fixing each sampled trajectory and re-scoring it under two opposing outcome conditions, one asserting correctness and the other incorrectness. Most affected tokens shift in the same direction under both conditions, with few sign reversals and substantial overlap in the induced optimization signals. Large shifts also concentrate on highly substitutable surface-form tokens, whereas tokens carrying problem-specific reasoning content are less sensitive. These findings show that privileged shifts fail to provide reliable answer-aligned directions, while their magnitudes primarily reflect counterfactual sensitivity rather than token-level learning value. Based on these observations, we propose Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve both the original credit budget and verifier-determined direction. On long-CoT mathematical reasoning benchmarks, CSCR consistently outperforms GRPO baseline with the same number of policy updates. Targeted ablations further corroborate our diagnosis: privilege-induced directions are unreliable, moderate downweighting is most effective, and stronger modulation destabilizes optimization.
Comments20 pages, 6 figures, 11 tables