发表机构
Aalto University; Shenzhen University of Advanced Technology; Zhejiang University(阿尔托大学; 深圳先进技术大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Nereus提出一种成本感知的自适应并行运行时,通过弹性模型单元和全局转换图,动态调整RL后训练的执行计划,显著降低延迟并提升吞吐量。
AI 中文摘要
大语言模型(LLM)的强化学习(RL)后训练在GPU集群上协调多个模型,涉及生成、推理和训练等环节。在运行过程中,多种因素可能发生变化,包括资源可用性、序列长度、内存压力以及阶段瓶颈。因此,最初合适的执行计划可能会随着时间推移变得缓慢甚至不可行。然而,调整一个模型共享GPU的任务面临重大挑战:判断新计划是否值得付出转换成本、复用任务的分布式状态,以及协调跨模型和跨阶段的GPU迁移。Nereus作为一个成本感知的运行时系统,针对这些挑战,将RL后训练任务自适应调整为高效的执行计划。其低开销控制器选择一个内存可行的全局计划,并通过针对运行中任务校准的成本模型来允许转换。为了估计并执行转换,Nereus将模型阶段(一个模型在一个阶段中)的每个副本的分布式状态表示为弹性模型单元(Elastic Model Unit)。然后,它利用全局转换图来排序这些单元的转换和GPU迁移。在基于真实数据构建的轨迹中,在线TP/PP自适应相对于初始固定TP/PP布局并配合DP扩展,将平均步骤延迟降低了27.7%。在达到1024个GPU的1000步运行中,六次转换仅占总运行时间的0.079%。在多样化的集群上,Nereus将端到端8B PPO吞吐量相对于OpenRLHF提升了2.14至7.27倍,相对于Verl提升了1.10至1.47倍。
英文摘要
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27$\times$ over OpenRLHF and by 1.10--1.47$\times$ over Verl across diverse clusters.