arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30528cs.LG

PAC:面向大语言模型多任务强化学习的进度增强优势课程

PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Yuanqiang Yu, Yanzhao Zheng, Zhentao Zhang, Tianze Xu, Chao Ma, Jihuai Zhu, Jiashun Liu, Xinle Deng, Baohua Dong, Hangcheng Zhu, Ruohui Huang

AI总结:

该研究针对LLM多任务RL后训练的任务分配问题,提出PAC方法结合优势与奖励增益信号,经多场景验证可提升样本效率与最终性能。

AI中文摘要:

强化学习(RL)被用于提升大语言模型(LLM)的推理能力,其训练数据涵盖异构任务。然而,大多数RL后训练流程依赖固定或人工设计的任务混合方案,即便任务的有用性会随训练推进而变化。在线课程方法常通过更新幅度定义可学习性,却忽略该更新是否转化为奖励增益,这可能将rollout预算分配给更新幅度大但无效的任务。我们提出PAC,即面向LLM多任务RL的进度增强优势课程,它结合两种任务级信号:优势衍生的可学习性(衡量任务能引发的策略更新幅度)和近期奖励增益(表明这些更新是否提升了任务性能)。贝叶斯汤普森采样控制器利用这些信号在GRPO训练期间分配跨任务的rollout。我们在两种场景下评估PAC:多级推理场景和多领域推理场景。PAC提升了样本效率和最终性能:它以更少的rollout步骤达到可比的验证分数,且在两种场景下均比随机采样和基于优势的课程基线实现更高的最终平均值。这些结果表明,联合追踪优势信号和实际奖励增益可为LLM后训练生成有效的在线课程。

英文摘要:

Reinforcement learning (RL) is used to improve the reasoning abilities of LLMs, while training data span heterogeneous tasks. However, most RL post-training pipelines rely on fixed or manually designed task mixtures, even though task usefulness changes as training progresses. Online curriculum methods often define learnability by update magnitude, ignoring whether the update translates into reward gains, which can misallocate rollout budget toward tasks with large but ineffective updates. We propose PAC, a Progress-Augmented Advantage Curriculum for multi-task RL of LLMs that combines two task-level signals: advantage-derived learnability, which measures the magnitude of the policy update a task can induce, and recent reward gains, which show whether those updates have improved task performance. A Bayesian Thompson Sampling controller uses these signals to allocate rollouts across tasks during GRPO training. We evaluate PAC under two settings: a multi-level reasoning setting and a multi-domain reasoning setting. PAC improves sample efficiency and final performance: it reaches comparable validation scores with fewer rollout steps and achieves higher final averages than random sampling and advantage-based curriculum baselines in both settings. These results show that jointly tracking advantage signals and actual reward gains yields an effective online curriculum for LLM post-training.

补充信息

↑