发表机构
vivo AI Lab; Department of Computer Science, Sun Yat-Sen University(vivo人工智能实验室; 中山大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型强化学习扩展的挑战,提出REGEN方法,通过回收重放记忆并运用离线强化学习算法训练通用模型,解耦训练过程,降低成本,在多任务上以低成本达类似准确率,还可能变革在线强化学习及扩展训练阶段。
AI 中文摘要
大规模在线强化学习是在大语言模型中引发包括长期推理和智能体工具使用等高级能力的主要手段。然而,在计算基础设施和成本方面,在大量感兴趣的任务领域继续扩展它仍然具有挑战性。多教师策略蒸馏有助于解耦强化学习阶段,但仍存在局限性。为此提出REGEN,通过回收重放记忆并采用离线强化学习算法训练通用模型,完全解耦了展开采样和反向训练过程,大幅降低训练成本,在多个任务上以更低成本达到多教师策略蒸馏的准确率,还可能将在线强化学习转变为数据合成过程并扩展到大规模训练后阶段。
英文摘要
Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs). However, continuing to scale it across vast task domains of interest remains challenging in both computational infrastructure and cost, especially when considering RL as merely a one-off learning stage. Recently, a widely used technique for distilling knowledge across various domains and training stages, multi-teacher on-policy distillation (MOPD), helps to decouple the RL stage, saving costs, while maintaining generality across vast domains. Nonetheless, similar to online RL, MOPD requires coupled inference and backward passes, which continues to limit its scalability and computational efficiency. To address these challenges, we propose REGEN: Replay-recycling for Expert-to-Generalist Distillation with Offline RL. Instead of distilling from multiple teacher models, REGEN trains a generalist by simply recycling the replay memory -- the free by-product of the teachers' specialized RL training -- and employing offline RL algorithms. REGEN completely decouples the rollout sampling from the backward training process and thus greatly reduces the training cost. Across mathematical reasoning, code generation, and instruction following, REGEN matches the accuracy of MOPD at substantially lower cost. It potentially turns online RL into a data synthesis process instead of a one-off learning stage, and can be extended to large-scale post-training without requiring heavy computational load. Code is available at https://github.com/yunjie-sysu/REGEN.