arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

半事实信用增强策略优化

Semifactual Credit-Augmented Policy Optimization

Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao, Qiaosheng Zhang, Yue Zhang

arXiv 2609.40360首次发表:更新:

发表机构

Zhejiang University; Westlake University; Shanghai Innovation Institute; Shanghai AI Laboratory(浙江大学; 西湖大学; 上海创新研究院; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RLVR中GRPO对token统一分配优势的局限,提出SCAPO,利用半事实稳定性进行细粒度信用分配,提升LLM推理准确率与泛化能力。

AI 中文摘要

基于可验证奖励的强化学习(RLVR)提升了大型语言模型(LLMs)的推理能力,但其预测仍对任务无关的提示特征敏感。我们通过保持底层问题及其答案不变的半事实提示干预来研究这种敏感性。我们的分析揭示了token级敏感性的显著变化,并表明在解码过程中抑制高漂移token候选可以在不更新模型权重的情况下提高推理准确性。这些发现凸显了组相对策略优化(GRPO)的一个局限性,即它为每个响应token分配相同的基于结果的优势,可能强化潜在虚假依赖而非有用推理。受此观察启发,我们引入了半事实信用增强策略优化(SCAPO),这是GRPO的一种因果启发变体,将半事实稳定性纳入token级信用分配。SCAPO在半事实干预下测量固定响应的token概率漂移,并使用归一化稳定性分数在早期训练中降低相对不稳定token的优势,同时不因稳定性本身给予额外信用。在Qwen3-4B-Base和Qwen3-1.7B-Base上,SCAPO在AIME 2024-2026上的准确率分别比GRPO提高了5.63和4.17个百分点。在两种模型规模下,SCAPO在大多数评估的数学基准和所有评估的分布外基准上均取得了最佳结果。这些结果表明,半事实稳定性通过RLVR中更细粒度的信用分配为改进推理和泛化提供了有效的训练信号。代码可在以下网址获取:此https URL。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑