arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DynaResize:用于分布式语言模型训练后运行时的GPU重新分配

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

Hanlin Du, Zhiyuan Yan, Yungang Bao, Sa wang

arXiv 2607.22614首次发表:更新:

发表机构

SKLP, Institute of Computing Technology, CAS; Bytedance(中国科学院计算技术研究所智能处理器研究中心; 字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于强化学习的语言模型训练后GPU资源分配问题,提出DynaResize运行时GPU重新分配系统,通过动态切换GPU平衡阶段执行时间,分解操作并去除关键路径非关键工作,实验显示可提高吞吐量、减少执行时间并隐藏部分开销。

AI 中文摘要

基于强化学习的语言模型训练后越来越多地将展开和训练分散到单独的GPU资源上,但静态GPU分区在长尾展开延迟下存在严重的流水线气泡。我们提出了DynaResize,这是一个运行时GPU重新分配系统,它在展开和训练之间动态切换GPU,以平衡阶段执行时间,而不改变强化学习语义。DynaResize将重新分配分解为细粒度操作,并通过通信器重用、有界状态暂存和基于滞后的重新分配,从关键路径中去除非启动关键工作。实验结果表明,DynaResize可以比最优静态配置将端到端吞吐量提高66.5%,并将总执行时间减少33%,同时隐藏27%的角色切换开销。

英文摘要

RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑