arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09207cs.DC

用于解耦和异步强化学习后训练的双向资源调度

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training

Zhiqiang Tan, Maoxin Wang, Sijie Wang, Yiming Yin, Qiang Wang, Xiaowen Chu, Shaohuai Shi

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对不同RL设置和工作负载下资源常空闲的问题,提出BiDiRL架构,通过开发热切换运行时、基于时间性能建模的规划器及双向调度器,减少资源空闲,在32-GPU测试平台上显著提高RL训练吞吐量。

中文摘要 AI 辅助

众所周知,通过在训练后期应用强化学习(RL)可以提高大语言模型(LLM)的推理能力。在标准RL迭代中,当前模型(策略)通过展开生成经验,然后在训练期间使用生成的数据更新策略。高性能RL框架如StreamRL和AReaL采用解耦架构和异步展开来更好地利用展开和训练资源,提高系统整体吞吐量。然而,在不同RL设置和变化的工作负载下,展开和训练资源仍常出现空闲期。本文提出BiDiRL,一种用于异步、解耦RL的混合时空复用架构,以减少资源空闲。首先开发热切换运行时,实现展开和训练资源间快速切换且开销可忽略。其次基于时间性能建模提出静态、调度感知规划器,选择利于热切换的资源分区,使展开和训练持续时间在粗粒度上大致平衡。最后在执行时引入双向调度器,通过细粒度资源切换进一步利用运行时空闲时间,让瓶颈阶段从另一资源池临时借用空闲资源。在两个32-GPU测试平台上,针对多种工作负载、数据集和模型,BiDiRL与包括veRL、AReaL和ROLL在内的RL系统相比,将RL训练吞吐量提高了1.94倍,且不影响收敛行为。

英文摘要

It is well established that the reasoning capabilities of large language models (LLMs) can be improved by applying reinforcement learning (RL) in a post-training stage. In a standard RL iteration, the current model (the policy) generates experience through rollouts, and the resulting data is then used to update the policy during training. High-performance RL frameworks such as StreamRL and AReaL employ a disaggregated architecture and asynchronous rollouts to better exploit both rollout and training resources, thereby increasing overall system throughput. Nonetheless, across varying RL setups (e.g., hardware configurations, model scales, staleness levels, and hyperparameters) and under changing workloads, it remains common for both rollout and training resources to experience idle periods. In this paper, we present BiDiRL, a hybrid time-space multiplexing architecture for asynchronous, disaggregated RL designed to reduce resource idleness. First, we develop a hot-switch runtime that enables rapid switching between rollout and training resources with negligible overhead. Second, we propose a static, scheduling-aware planner based on time-performance modeling that chooses a hot-switch-friendly resource partition, so that rollout and training durations are roughly balanced at a coarse level. Third, at execution time, we introduce a bidirectional scheduler that further exploits runtime bubbles through fine-grained resource switching, allowing the bottleneck stage to temporarily borrow idle resources from the other pool. Across a wide range of workloads, datasets, and models on two 32-GPU testbeds, BiDiRL increases RL training throughput by up to 1.94x compared with RL systems including veRL, AReaL, and ROLL, without affecting convergence behavior.

补充信息

↑