SCOPE-RL:成功前后优化推理路径
SCOPE-RL: Optimizing Reasoning Paths Before and After Success
- Baidu Inc.(百度公司)
- Shandong University(山东大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对可验证奖励强化学习中推理路径反馈不足问题,提出SCOPE-RL框架,分两阶段优化,成功前添加奖励,成功后细化轨迹,并经评估协议验证,相比仅结果的GRPO提升准确率、减少推理令牌,且与其他方法互补。
AI中文摘要:
可验证奖励的强化学习(RLVR)使用稀疏可验证的最终答案奖励来优化语言模型。这种稀疏锚点能可靠验证轨迹是否成功,但对产生成功的推理路径无直接反馈。成功前,难题的前期进展无奖励信号;成功后,结果奖励无法区分良好组织的正确轨迹与冗余或局部有缺陷的轨迹。我们引入SCOPE-RL,分两阶段强化锚点并保留GRPO更新:成功前,自适应支架式强化学习在答案隐藏子问题链上添加前缀分解可验证奖励;成功后,质量感知过程强化学习应用正确性门控过程形状奖励来优化正确轨迹。专家验证的步骤质量评估协议评估有用步骤密度、错误定位和令牌效率。在Qwen3-8B-Instruct上训练,SCOPE-RL比仅基于结果的GRPO平均准确率提高11.2个百分点,推理令牌减少27.1%,在GSPO和Qwen3-0.6B-Instruct上也有提升,表明奖励信号强化与策略更新级RLVR进展互补。
英文摘要:
Reinforcement learning with verifiable rewards (RLVR) optimizes LLMs using sparse verifiable final-answer rewards. This sparse anchor reliably verifies whether a trajectory succeeds but provides no direct feedback on the reasoning path that produced it. Before success, prerequisite progress on hard problems receives no reward signal; after success, outcome rewards cannot distinguish well-organized correct trajectories from redundant or locally flawed ones. We introduce SCOPE-RL (Scaffolded Chain Optimization with Process Efficiency), a two-stage framework that densifies this anchor while retaining the GRPO update: Adaptive Scaffolded RL adds prefix-decomposed verifiable rewards on answer-hidden sub-question chains before success, and Quality-Aware Process RL applies correctness-gated process-shape rewards to refine correct trajectories after success. An expert-validated Step-Quality Evaluation Protocol evaluates useful-step density, error localization, and token efficiency beyond final-answer accuracy. On Qwen3-8B-Instruct trained on DAPO-Math and Big-Math, SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO; the gains hold under GSPO and on Qwen3-0.6B-Instruct, indicating that reward-signal densification is complementary to policy-update-level RLVR advances. Code and data are available at https://github.com/tokencraft-lab/SCOPE-RL.