发表机构
Beijing University of Posts Telecommunications; StepFun(北京邮电大学; 阶跃星辰)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CadenceRL通过结构性的工作负载重塑(节奏控制与集中)解决异构回放池中异步RL后训练的调度难题,显著提升解码吞吐量并降低轨迹延迟。
AI 中文摘要
强化学习(RL)后训练日益依赖于长视野、多轮次的回放(rollout)。当后训练任务超出单个集群的规模时,跨集群组装的回放池会引入硬件异构性。回放调度必须服务于两个利益相关方:硬件需要高聚合解码吞吐量,而每个轨迹需要快速完成。这种张力源于自回归解码的内存带宽受限特性。大的活动批次摊销权重读取以实现高吞吐量,但给每个轨迹留下的带宽份额更小,完成时间更长。因此,调度目标是专业化,让不同的工作节点扮演不同的角色。异构硬件进一步实现了这种专业化。高带宽加速器有利于长上下文工作,而成本高效的加速器则维持大批次。工作负载的演变使得这种专业化难以维持,动态重新分配面临循环依赖,因为移动的收益取决于后续的放置决策。我们提出了CadenceRL,它通过结构性的工作负载重塑而非逐移动收益估计来绕过这种依赖。节奏控制(Pacing)用较短的轨迹替换长上下文轨迹,提供了一种结构性的正向变换,以维持大批次活动批次实现高吞吐量。当累积的陈旧性要求更快的完成时,集中(Concentration)将残余的长上下文尾部引导到高亲和力的工作节点上。延迟绑定的KV准备阶段在目的地选择之前累积前缀。在异构回放池上,CadenceRL将解码吞吐量提高了高达48%,并将P95轨迹延迟降低了高达64%。添加高带宽加速器可减少尾部延迟,而添加成本高效的加速器可提高吞吐量,无需手动路由配置。
英文摘要
Reinforcement learning (RL) post-training increasingly relies on long-horizon, multi-turn rollouts. As post-training jobs outgrow a single cluster, rollout pools assembled across clusters introduce hardware heterogeneity. Rollout scheduling must serve two stakeholders: the hardware needs high aggregate decode throughput, while each trajectory needs to finish quickly. The tension arises from the memory-bandwidth-bound nature of autoregressive decoding. A large active batch amortizes weight reads for high throughput but leaves each trajectory a smaller bandwidth share and a longer completion time. The scheduling objective is therefore specialization, letting different workers serve different roles. Heterogeneous hardware further enables this specialization. High-bandwidth accelerators favor long-context work, while cost-efficient accelerators sustain large batches. Workload evolution makes this specialization difficult to sustain, and dynamic reassignment faces a circular dependency because a move's benefit depends on subsequent placement decisions. We present CadenceRL, which bypasses this dependency through structural workload reshaping rather than per-move benefit estimation. Pacing replaces long-context trajectories with shorter ones, providing a structurally positive transformation that sustains large active batches for high throughput. When accumulated staleness demands faster completion, concentration directs the residual long-context tail onto high-affinity workers. Late-bound KV preparation stages accumulated prefixes before a destination is selected. On heterogeneous rollout pools, CadenceRL improves decode throughput by up to 48% and reduces P95 trajectory latency by up to 64%. Adding high-bandwidth accelerators reduces tail latency, while adding cost-efficient accelerators increases throughput, without manual routing configuration.