Conduit:面向分布式强化学习的经验数据平面
Conduit: An Experience Data Plane for Distributed Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
Conduit提出经验数据平面(EDP),将经验管理作为显式系统优化问题,通过容量受限、带宽感知的放置和延迟感知调度,显著降低经验路径延迟,可扩展至1024个GPU并保持收敛。
中文摘要 AI 辅助
分布式强化学习(RL)通过围绕经验缓冲区并行化参与者和学习器来扩展训练规模。然而,随着强化学习工作负载的增长,缓冲区不再仅仅是一个回放队列:它是大容量、延迟敏感的经验路径的存储基础,每次迭代都要遍历该路径以移动、转换、采样和批处理经验,之后学习器更新才能开始。现有的强化学习系统将此路径嵌入框架控制流中,或将其作为请求驱动的缓冲区服务暴露出来,导致经验放置固定不变,经验路径工作难以作为运行时级别的优化目标独立调度。我们提出了Conduit,一个与框架无关的运行时,将强化学习经验管理明确作为一个系统优化问题。其核心是经验数据平面(EDP),一种运行时抽象,通过将经验摄取、经验放置和经验交付作为显式控制点暴露,将强化学习经验处理语义与框架特定的执行逻辑分离。基于EDP,Conduit引入了容量受限、带宽感知的放置策略,该策略在异构互连和设备内存约束下,将经验状态分布到CPU/GPU内存层级和节点上;以及延迟感知的调度策略,该策略控制经验路径处理何时运行,以减少暴露的经验路径延迟,同时保持强化学习语义。与RLlib集成且不改变其框架执行逻辑,Conduit将暴露的经验路径延迟降低高达97%,端到端迭代延迟降低高达38%,可扩展到1,024个GPU,并保持收敛性。
英文摘要
Distributed reinforcement learning (RL) scales training by parallelizing actors and learners around an Experience Buffer. As RL workloads grow, however, the buffer becomes more than a replay queue: it is the storage substrate of a large-capacity, latency-critical experience path that every iteration traverses to move, transform, sample, and batch experiences before learner updates can begin. Existing RL systems embed this path inside framework control flow or expose it as a request-driven buffer service, leaving experience placement fixed and experience-path work difficult to schedule independently as a runtime-level optimization target. We present Conduit, a framework-agnostic runtime that exposes RL experience management as an explicit systems optimization problem. At its core is the Experience Data Plane (EDP), a runtime abstraction that separates RL experience-handling semantics from framework-specific execution logic by exposing experience ingestion, experience placement, and experience delivery as explicit control points. Built on EDP, Conduit introduces capacity-constrained, bandwidth-aware placement, which distributes experience state across CPU/GPU memory tiers and nodes under heterogeneous interconnect and device-memory constraints, and latency-aware scheduling, which controls when experience-path handling runs to reduce exposed experience-path latency while preserving RL semantics. Integrated with RLlib without changing its framework execution logic, Conduit reduces exposed experience-path latency by up to 97% and end-to-end iteration latency by up to 38%, scales to 1,024 GPUs, and preserves convergence.
发表机构
- Aalto University(阿尔托大学)
- Shenzhen University of Advanced Technology(深圳先进技术大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。