发表机构
Nanjing University; Institute for AI Industry Research (AIR), Tsinghua University; Nanjing University of Posts and Telecommunications; University of Science and Technology of China(南京大学; 清华大学智能产业研究院(AIR); 南京邮电大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对具身强化学习中资源利用不均和同步停滞问题,提出EBRL异步训练系统,通过异步流水线调度和细粒度资源管理,实现1.30-3.47倍吞吐量和2.5倍收敛速度提升。
AI 中文摘要
具身强化学习通过环境模拟、动作生成和模型更新的流水线来提升模型能力。这些阶段对CPU和GPU的需求各异,使得高效资源利用变得困难。近期系统为提升效率将展开(模拟与生成)与训练重叠,但独占GPU分配和展开中的同步屏障仍造成大量硬件资源浪费。本文提出EBRL,一个采用两项核心技术的异步具身强化学习训练系统。异步流水线调度器将展开与训练重叠,跨环境组对模拟与生成进行流水线处理,并独立执行每个环境,从而消除同步停滞。细粒度资源管理器汇集CPU核心和GPU流式多处理器,并利用阶段画像和运行时反馈调整资源配额和批大小,以满足各阶段间变化的需求。我们在RLinf上实现EBRL,并在异构GPU测试台上用四种具身策略和四个模拟基准进行评估。实验表明,与最先进的具身强化学习系统相比,EBRL实现了1.30-3.47倍的端到端展开吞吐量和2.5倍的训练收敛速度。
英文摘要
Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. These stages show heterogeneous CPU and GPU demands, making efficient resource utilization difficult. Recent systems overlap rollout (simulation and generation) with training for efficiency, but exclusive GPU allocation and synchronized barrier in rollout still leave substantial hardware resource waste. In this paper, we present EBRL, an asynchronous embodied RL training system with two core techniques. The asynchronous pipelined scheduler overlaps rollout and training, pipelines simulation and generation across environment groups, and carries out each environment independently, eliminating synchronization stalls. The fine-grained resource manager pools CPU cores and GPU streaming multiprocessors, and uses stage profiles and runtime feedback to adjust resource quotas and batch sizes to meet the shifting demands among stages. We implement EBRL on RLinf and evaluate it with four embodied policies and four simulation benchmarks across heterogeneous GPU testbeds. Experiments show that EBRL achieves 1.30-3.47 times the end-to-end rollout throughput and 2.5 times of training convergency compared to the SOTA embodied RL systems.