arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理性引导的策略优化:利用自适应理性脚手架学习推理

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei

arXiv 2610.07342首次发表:更新:

发表机构

New York University; University at Buffalo; Monash University(纽约大学; 布法罗大学; 莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对强化学习奖励稀疏问题,提出理性引导策略优化(RGPO),自适应利用真实理性作为临时脚手架,保留探索自由,在语言和视觉-语言推理任务上持续超越基线。

AI 中文摘要

在策略强化学习已成为提升大型语言模型推理能力的核心范式。然而,其有效性常常受到奖励稀疏性的限制:当模型无法为难题发现正确的轨迹时,优化过程接收到的有用信号很少,可能会停滞不前。现有方法通过引入离策略演示、专家轨迹或模型生成的解决方案来缓解这一问题,但这些方法通常要求辅助数据与强化学习任务的格式相匹配,往往依赖于从更强模型中进行拒绝采样以获得合适的训练轨迹。我们提出了理性引导的策略优化(RGPO),这是一个根据模型当前能力自适应地利用真实理性信息、同时保留其探索自由的框架。RGPO 并非将参考解决方案视为固定的模仿目标,而是将其用作临时脚手架:理性信息帮助模型生成改进的响应,之后仅将获得更高奖励的、模型生成的解决方案转移回原始的无引导设置。这种设计使得训练能够利用可用的真实信息,而无需离策略数据遵循与强化学习任务相同的格式。在纯语言和视觉-语言推理设置中,RGPO 均持续优于 RLVR 基线,消融研究表明自适应理性引导是这些收益的关键贡献因素。这些结果表明,RGPO 提供了一种实用且通用的方法,用于减少奖励稀疏性、稳定强化学习,并在纯文本和多模态模型中提升推理性能。

英文摘要

On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptively leverages ground-truth rationale information according to the model's current capability while preserving its freedom to explore. Rather than treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model generate improved responses, after which only higher-reward, model-generated solutions are transferred back to the original unguided setting. This design allows training to exploit available ground-truth information without requiring off-policy data to follow the same format as the RL task. Across both language-only and vision-language reasoning settings, RGPO consistently improves performance over RLVR baselines, and ablation studies show that adaptive rationale guidance is a key contributor to these gains. These results suggest that RGPO offers a practical and general approach for reducing reward sparsity, stabilizing reinforcement learning, and improving reasoning performance in both text-only and multimodal models.

CommentsNeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑