发表机构
Aalto University; NVIDIA(阿尔托大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NeMo-DCR通过位精确增量压缩重配,仅传输变化权重,使万亿参数智能体RL的重配速度提升12-40倍,1T模型重配从87.5分钟降至150秒。
AI 中文摘要
智能体强化学习(RL)将训练与采样(rollout)分离,因此每次策略更新都必须在下一批次之前到达采样集群。在两个AWS区域之间传输完整的1T检查点以进行这种权重同步(重配)需要87.5分钟。对BF16训练的测量表明,每一步约有1%的权重会改变其存储值。近期系统利用了这一稀疏性,但在放置、精确性或效率方面存在不足:它们重新实现放置规则、组装完整张量、以算术方式重建值或使用跨集群集合通信,且没有系统能从重配中途失败中完全恢复。我们提出NeMo-DCR(增量压缩重配),它仅发送变化但保持位精确:接收方获得与密集重配相同的参数和缓冲区位。对于放置,固定仿射映射将训练分片中的变化投影到检查点的规范坐标,残差转换覆盖其他变化,服务运行时的原生加载器将所有变化放置在接收方存储中。对于精确性,可压缩的XOR掩码携带仿射变化,其投影和加载器保留存储位,覆盖写入携带其他变化。接收方就地应用两者,重试覆盖部分写入,联合提交将策略绑定到用于下一个增量的基线。对于效率,对象存储或中继树在增量构建期间流式传输负载,无需跨集群集合通信。即使在3%和5%的变化率下,NeMo-DCR对30B-1T模型的重配比仅传输的完整检查点参考快12-40倍。在3%变化率下,1T中继树重配耗时150秒而非87.5分钟,使重配在万亿参数规模的跨集群智能体RL中变得实用。
英文摘要
Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint's canonical coordinates, residual conversion covers the other changes, and the serving runtime's native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40$\times$ faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.
CommentsThe code is open-sourced in NVIDIA NeMo RL PR #2444 at https://github.com/NVIDIA-NeMo/RL/pull/2444