WeightBridge:一种用于强化学习的高效权重传输库
WeightBridge: An Efficient Weight Transfer Library for Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
WeightBridge是一个高效灵活的权重传输库,通过自动提取布局对应关系并规划无冗余、负载均衡的传输,在多种RL配置下将GPU停顿时间平均减少高达42倍,并易于集成。
中文摘要 AI 辅助
权重传输——即从训练器向回放生成器传播更新后的参数——正成为面向大语言模型的强化学习(RL)系统中的重要性能瓶颈。核心挑战在于,在不牺牲效率的前提下,支持现代RL工作负载中多样化的训练器与回放布局以及同步需求。现有解决方案在某些配置下高效,但在其他配置下表现不佳或缺乏支持。我们提出了WeightBridge,一个灵活、高效的权重传输库,旨在跨多种RL配置提供高性能。WeightBridge首先自动提取训练器与回放权重布局之间的对应关系,然后规划并执行无冗余且负载均衡的权重传输。它暴露了一个小而通用的API,同时协调跨多种同步模式的工作节点。在涵盖不同模型、并行化布局和同步模式的多种配置中,WeightBridge相比最先进的开源RL框架,平均将GPU停顿时间减少了高达42倍,并在所有设置中实现了高性能。一个编码智能体能够在无需人工指导的情况下将WeightBridge集成到两个不同的RL框架中,展示了其API的通用性和易用性。
英文摘要
Weight transfer - the propagation of updated parameters from trainers to rollout generators - is becoming an important performance bottleneck in reinforcement learning (RL) systems for LLMs. The central challenge is supporting the diverse trainer and rollout layouts and synchronization requirements of modern RL workloads without sacrificing efficiency. Existing solutions are efficient under some configurations but perform poorly or lack support under others. We present WeightBridge, a flexible, efficient weight-transfer library designed to deliver high performance across diverse RL configurations. WeightBridge first automatically extracts the correspondence between trainer and rollout weight layouts, then plans and executes redundancy-free and load-balanced weight transfer. It exposes a small, general API while coordinating workers across diverse synchronization modes. Across configurations spanning different models, parallelization layouts, and synchronization modes, WeightBridge reduces average GPU stall time by up to 42$\times$ over the state-of-the-art open-source RL framework and achieves high performance in all settings. A coding agent was able to integrate WeightBridge into two different RL frameworks without manual guidance, demonstrating the generality and ease of use of its APIs.
发表机构
- FAIR at Meta(Meta人工智能研究院)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。