发表机构
Peking University; Baidu Inc.; Kuaishou Inc(北京大学; 百度公司; 快手公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对强化学习中稀疏奖励致令牌级信用分配难的问题,提出基于模式局部替代熵的ACPO框架,通过不对称调制策略更新改进信用分配,实验表明其优于多种强化学习基线。
AI 中文摘要
强化学习提升了大语言模型推理能力,但稀疏结果奖励使令牌级信用分配困难。现有方法存在问题,本文提出ACPO,基于模式局部替代熵的令牌级信用分配框架,不对称调制策略更新,实验显示其在数学推理和编码基准测试中优于强基线。
英文摘要
Outcome-supervised reinforcement learning scales to verifiable reasoning tasks, but trajectory-level rewards assign the same outcome signal to all sampled tokens, overlooking their unequal contributions to the reasoning process. Entropy provides a natural indicator of the model's decision state, yet using it for token-level credit assignment presents two key challenges: long-tail probabilities in large vocabularies corrupt both entropy values and gradients, and uncertainty carries distinct semantics across positive- and non-positive-advantage trajectories. We propose Asymmetric Credit Policy Optimization (ACPO), which replaces global entropy with the complement of the top-token probability as a mode-local proxy. Guided by gradient analysis, ACPO incorporates mismatch routing and saturation correction to shape policy updates into the desired asymmetric form, emphasizing uncertain decisions on positive trajectories while penalizing confident regions on failed ones. Theoretically, ACPO locally preserves the advantage direction while bounding surrogate error. Experiments on mathematical and coding reasoning benchmarks, including AIME 2025 and HumanEval Pro, show that ACPO consistently outperforms both entropy-aware methods (e.g., 80/20, GTPO) and strong outcome-supervised RL baselines (e.g., DAPO, SAPO).