arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AReaL-DTE:面向在线智能体强化学习的稀疏权重迁移技术

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning

Yingqi Peng, Jiawei Zhang, Wenhao Zhou, Ruida Xu, Ran Yan, Wei Dong, Yi Gao, Zhiqiang Ding, Tongkai Yang, Binhang Yuan

arXiv 2608.00455首次发表:更新:

AI 中文总结

本研究提出面向在线智能体强化学习的 AReaL-DTE,通过无快照增量传输引擎实现稀疏权重迁移,在跨集群及集群内场景下大幅提升速度并降低内存占用。

AI 中文摘要

采用微服务实现的在线智能体强化学习将策略训练与 rollout 生成分离,提升了可扩展性与模块化程度,但可能使频繁的策略权重同步成为关键的系统开销。共享存储自然地跨集群连接这些服务,而普通的密集策略权重同步可能会产生模型构建、传输及应用的成本。稀疏同步减少了传输数据量,但面向 checkpoint 的方法仍会保留先前模型并实例化完整中间结果,以桥接异构的训练与推理布局。我们提出 AReaL-DTE,一种无快照的 Delta Transfer Engine(增量传输引擎),它将推理可见的权重稀疏性转化为端到端的系统效率。在我们评估的工作负载中,连续策略版本之间仅有不到 2% 的 BF16 权重元素发生变化。AReaL-DTE 通过反转 AdamW 更新按需重构被覆盖的权重,通过与转换器对齐的 BF16 变化检测传输重构及当前参数,并将变化元素直接重映射到接收方本地坐标。AReaL-DTE 支持通过共享存储跨集群实现清单提交的稀疏传输,以及在集群内采用死锁安全的两轮协议,随后直接应用于推理分片。我们在 Qwen3-8B 和 Qwen3-30B-A3B 上,针对四项在线 RL 工作负载评估了 AReaL-DTE。AReaL-DTE 在跨集群场景下,相比 ByteCheckpoint 实现了最高 19.9 倍的加速,相比 PULSE 实现了最高 3.2 倍的加速;在集群内场景下,分别实现了最高 7.6 倍和 7.4 倍的加速。在集群内 Qwen3-30B-A3B 实验中,它将峰值 GPU 内存降低了约 41%,峰值 CPU 内存至少降低了 87%。

英文摘要

Online agentic reinforcement learning implemented with micro-services separates policy training from rollout generation, improving scalability and modularity while potentially making frequent policy-weight synchronization a critical systems overhead. Shared storage naturally connects these services across clusters, but vanilla dense policy weight synchronization could incur model-scale construction, transfer, and application costs. Sparse synchronization reduces transferred data, yet checkpoint-oriented approaches can still retain a previous model and materialize complete intermediates to bridge heterogeneous training and inference layouts. We present AReaL-DTE, a snapshot-free Delta Transfer Engine that translates inference-visible weight sparsity into end-to-end system efficiency. Across our evaluated workloads, fewer than 2% of BF16 weight elements change between consecutive policy versions. AReaL-DTE reconstructs overwritten weights on demand by inverting AdamW updates, streams reconstructed and current parameters through converter-aligned BF16 change detection, and remaps changed elements directly into receiver-local coordinates. AReaL-DTE supports manifest-committed sparse transfer through shared storage across clusters and a deadlock-safe two-round protocol within a cluster, followed by direct application to inference shards. We evaluate AReaL-DTE on Qwen3-8B and Qwen3-30B-A3B across four online RL workloads. AReaL-DTE achieves speedups of up to 19.9x over ByteCheckpoint and 3.2x over PULSE across clusters, and up to 7.6x and 7.4x, respectively, within a cluster. In the same-cluster Qwen3-30B-A3B experiments, it reduces peak GPU memory by approximately 41% and peak CPU memory by at least 87%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑