arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CVPO:通过值方差自适应与动态课程学习增强大语言模型的强化学习推理能力

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning

Ziqi Jia, Yalu Ouyang, Bo Pang, Panpan Li, Hangfei Xu, Shengzhao Wen, Shiyong Li, Yanpeng Wang

arXiv 2608.03068首次发表:更新:

AI 中文总结

针对LLM强化学习推理中反馈精度不足与问题难度漂移问题,提出CVPO方法,通过值方差自适应与动态课程学习优化,在数学任务上性能优于VAPO等基线。

AI 中文摘要

强化学习(RL)已成为增强大语言模型(LLM)推理能力的有效方法。然而,现有方法存在生成答案轨迹反馈精度不足的问题,且会出现问题难度漂移现象。为应对这些挑战,我们提出CVPO——基于课程引导的值方差策略优化(Curriculum-guided Value-Variance Policy Optimization)。在响应轨迹层面,我们发现token级别的值方差与探索强度相关,理论分析表明该方差限定了策略更新幅度,随后我们利用估计的轨迹值方差量化生成过程中的内在随机性,基于此为不同奖励类型设计了感知方差的优势调整机制。在问题层面,我们引入了动态课程加权方法,该方法可适配问题难度,助力模型在每个训练阶段聚焦于与其当前能力匹配的任务。实验结果表明,我们的方法优于VAPO等强大的基于值的基线方法,在各类数学任务中,该方法实现了更优的性能与更强的探索能力,使语言模型的推理更准确、更鲁棒。

英文摘要

Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, existing methods suffer from insufficient precision in feedback on generated answer trajectories and exhibit the phenomenon of problem difficulty drift. To address these challenges, we propose CVPO - Curriculum-guided Value-Variance Policy Optimization. At the response trajectory level, we find that token-level value-variance correlates with exploration intensity. Our theoretical analysis shows this variance bounds policy update magnitude. We then use the estimated trajectory value-variance to quantify the intrinsic randomness in generation. Based on this, we design a variance-aware advantage adjustment mechanism for different reward types. At the question level, we introduce a dynamic curriculum weighting method that adapts to question difficulty. This helps the model focus on tasks matched to its current ability during each training stage. Experimental results show our method outperforms strong value-based baselines like VAPO. It achieves better performance and stronger exploration, enabling more accurate and robust reasoning in language models across various math tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑