发表机构
SKLP, Institute of Computing Technology, CAS; Bytedance(中国科学院计算技术研究所智能处理器研究中心; 字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于强化学习的语言模型训练后GPU资源分配问题,提出DynaResize运行时GPU重新分配系统,通过动态切换GPU平衡阶段执行时间,分解操作并去除关键路径非关键工作,实验显示可提高吞吐量、减少执行时间并隐藏部分开销。
AI 中文摘要
基于强化学习的语言模型训练后越来越多地将展开和训练分散到单独的GPU资源上,但静态GPU分区在长尾展开延迟下存在严重的流水线气泡。我们提出了DynaResize,这是一个运行时GPU重新分配系统,它在展开和训练之间动态切换GPU,以平衡阶段执行时间,而不改变强化学习语义。DynaResize将重新分配分解为细粒度操作,并通过通信器重用、有界状态暂存和基于滞后的重新分配,从关键路径中去除非启动关键工作。实验结果表明,DynaResize可以比最优静态配置将端到端吞吐量提高66.5%,并将总执行时间减少33%,同时隐藏27%的角色切换开销。
英文摘要
RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.