发表机构
University of Virginia(弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PACE方法,通过分解轨迹陈旧性为生成与等待两类,并利用池感知的拒绝预算和有效陈旧性评分,在不牺牲异步效率的前提下控制策略滞后,显著提升LLM后训练中异步RL的验证性能并节省GPU时间。
AI 中文摘要
完全异步的强化学习(RL)通过将轨迹生成与策略优化重叠,提高了大型语言模型后训练中的资源利用率,但同时也引入了策略滞后,因为轨迹在生成和排队时训练器仍在继续更新。我们研究了这种滞后如何在轨迹的生命周期内累积,以及如何在不牺牲异步执行带来的挂钟时间优势的情况下对其进行控制。我们将轨迹陈旧性分解为生成陈旧性(在轨迹完成前累积)和等待陈旧性(在完成的轨迹进入池后累积)。基于这一分解,我们引入了PACE(池感知的有效陈旧性控制)。PACE将多余的池占用转化为自适应拒绝预算,并使用结合等待陈旧性与前缀感知生成陈旧性的有效陈旧性评分对完成的轨迹进行排序。这避免了仅因轨迹跨越多个策略版本而惩罚长轨迹或中断的轨迹。在单轮数学推理中,PACE在相同的挂钟时间预算下,将六个基准的平均验证准确率比未过滤的异步RL提高了18.7%,并以减少47.1%的GPU时间达到了与同步RL相当的性能。PACE还在多轮工具集成推理中提高了验证性能,优于同步和未过滤的异步RL。进一步的混合专家模型实验和替代RL算法实验支持了其在不同模型架构和训练算法中的适用性。
英文摘要
Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess pool occupancy into an adaptive rejection budget and ranks completed trajectories using an effective-staleness score that combines Waiting Staleness with prefix-aware Generation Staleness. This avoids penalizing long or interrupted rollouts solely because they span multiple policy versions. In single-turn mathematical reasoning, PACE improves the six-benchmark average validation accuracy by 18.7\% over unfiltered asynchronous RL at the same wall-clock budget and matches synchronous RL performance with 47.1\% less GPU time. PACE also improves validation performance in multi-turn tool-integrated reasoning, outperforming both synchronous and unfiltered asynchronous RL. Further experiments with the mixture-of-experts model and an alternative RL algorithm support its applicability across model architectures and training algorithms.