发表机构
Tsinghua University; Renmin University of China; Institute of Information Engineering, CAS; USTC; The Australian National University; Peking University; University of Macau(清华大学; 中国人民大学; 中国科学院信息工程研究所; 中国科学技术大学; 澳大利亚国立大学; 北京大学; 澳门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出梯度对齐奖励(GAR)方法,在策略梯度空间中生成密集推理感知奖励,提升大语言模型在数学等基准上的推理性能,且无需领域特定数据即可迁移。
AI 中文摘要
可验证奖励强化学习(RLVR)驱动大语言模型的思维链推理,但其二元结果奖励无法区分正确轨迹。现有替代的密集奖励,从表面启发式到过程奖励模型,要么忽略训练语料中已有的专家解决方案,要么需要昂贵的离线标注。我们提出梯度对齐奖励(GAR),其在策略自身的梯度空间中运行:通过输出投影层的截断反向传播为每次rollout提取紧凑梯度向量,与专家锚点梯度的余弦相似度产生密集的、感知推理的奖励,且墙钟开销低于9%。我们证明该余弦可分解为预测误差和激活模式因子,为对齐信号的测量提供具体表征。在Qwen3-4B和Qwen3-8B上,GAR在竞赛级数学基准上始终优于GRPO及其他基线,且无需领域特定数据即可迁移至GPQA Diamond和MMLU-Pro。代码和数据可在此URL获取。
英文摘要
Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, either ignore the expert solutions already present in training corpora or require expensive offline annotation. We propose Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock overhead. We prove that this cosine admits a multiplicative decomposition into prediction-error and activation-pattern factors, providing a concrete characterization of what the alignment signal measures. On Qwen3-4B and Qwen3-8B, GAR consistently improves over GRPO and other baselines on competition-level math benchmarks and transfers to GPQA Diamond and MMLU-Pro without domain-specific data. Code and data are available at https://github.com/LQgdwind/GAR.