发表机构
Trendinsight Lab / UC San Diego(趋势洞察实验室/加州大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对高分辨率视频扩散模型内存消耗大问题,提出MegaSlide-DiT,通过让GPU不持有模型状态及用3D-DSA取代全局注意力,降低内存和计算复杂度,实现105B DiT在单H200 GPU上的适配,为大规模视频扩散模型全参数适配提供实用路径。
AI 中文摘要
基于扩散变换器(DiTs)构建的高分辨率视频扩散模型具有高保真度,但会迅速耗尽单个工作站的内存预算。参数超过1000亿的DiT轻易就需要超过1TB的持久状态,而朴素的时空自注意力在序列长度上呈二次方增长。这两个内存限制阻碍了研究人员在没有大型GPU集群的情况下适配大规模生成模型。我们从系统角度重新审视这个问题,引入了MegaSlide-DiT,展示了如何在具有1.5TB主机内存的单个H200 GPU上适配预训练的105B DiT。关键在于GPU无需拥有模型状态,所有持久权重等保留在主机内存,仅按需将临时分片传输到GPU。同时,用3D可变形滑动注意力(3D-DSA)取代二次方的全局注意力,它将内存和计算复杂度降至序列长度的线性。我们报告了详细的内存计算、执行轨迹和评估结果来证实设计。MegaSlide-DiT并非在单个GPU上从头训练105B模型,也未神奇解决带宽限制,而是为在高端工作站上对大规模视频扩散模型进行全参数适配提供了实用路径。
英文摘要
High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a single workstation. A 100 billion-plus parameter DiT easily requires over a terabyte of persistent state, while naive spatiotemporal self-attention grows quadratically in sequence length. These two walls -- parameter memory and activation memory -- prevent researchers from adapting massive generative models without large GPU clusters. We revisit this problem from a systems perspective and introduce MegaSlide-DiT, a prototype that demonstrates how a pre-trained 105B DiT can be adapted on a single H200 GPU with 1.5 TB of host RAM. Our key insight is that the GPU need not own the model state: all persistent weights, master weights and optimizer moments remain in host memory, while only transient shards are streamed to the GPU on demand. Simultaneously, we replace quadratic global attention with 3D Deformable Slide Attention (3D-DSA), a motion-adaptive local attention operator that reduces both memory and computational complexity to linear in the sequence length. We report detailed memory accounting, execution traces and evaluation results to substantiate our design. MegaSlide-DiT does not claim to train a 105B model from scratch on a single GPU, nor does it magically solve bandwidth limits; rather, it offers a pragmatic path for full-parameter adaptation of massive video diffusion models on high-end workstations.