发表机构
USTC; HKUST; Alibaba Group(中国科学技术大学; 香港科技大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体强化学习中多轮 rollout 成本高、训练 GPU 闲置的问题,提出 PEARL 系统,通过弹性资源协调、空闲训练 GPU 复用和自适应预填充-解码配置,显著提升吞吐量。
AI 中文摘要
多轮 rollout 主导了智能体强化学习(RL)的成本。异步执行和弹性 GPU 资源可以加速这一阶段,但增加 rollout 副本会产生收益递减,而训练 GPU 在更新之间仍处于空闲状态。我们观察到,有效的资源利用还取决于预填充-解码(PD)配置。共置与分离的选择以及最优 PD 比例均随工作负载而变化,这使得资源扩展与 PD 配置相互依赖。利用这一机会需要选择有效的配置,并在瞬态资源可用窗口内实现其收益,尽管存在重新配置成本。我们提出了 PEARL,一个异步智能体 RL 系统,它协调外部资源弹性、空闲训练 GPU 的临时复用以及自适应 PD 执行。PEARL 维护统一的 GPU-工作器-角色状态,并使用运行时配置文件预测 rollout 批次的完成时间,同时考虑环境引起的解码并发度降低。它在当前 GPU 预算下选择 PD 模式和比例,并将每个决策转化为最小化工作器和角色变更的增量转换计划。成本感知的切换和借用策略抑制预期收益不足的转换,同时确保训练 GPU 的及时归还。我们的评估表明,在不同 LLM 上,PEARL 的吞吐量是固定资源 ROLL 的 2.17-2.79 倍。与 RLBoost+ 相比,Qwen3-8B 的吞吐量提升最高约 26.9%,Qwen3-30B-A3B 提升约 36.3%。
英文摘要
Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill--decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves $2.17$--$2.79\times$ the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9\% for Qwen3-8B and 36.3\% for Qwen3-30B-A3B.