发表机构
NVIDIA Research, Efficient AI Team & Singapore Lab(NVIDIA Research,高效AI团队与新加坡实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MiniMax-H3视频扩散模型在云与边缘设备的推理瓶颈,提出全栈优化流水线,含跨分辨率调度与递归自我改进内核搜索,实现30倍加速和20%显存降低。
AI 中文摘要
视频扩散模型正在快速扩展并展现出更强的生成能力。在这些最新进展中,MiniMax-H3作为一个高能力、生产级的开源模型脱颖而出。然而,其330亿参数和逐步迭代去噪过程引入了巨大的计算开销。因此,其实际生产受到云部署(如NVIDIA-GB200)中生成延迟的阻碍,同时严格的显存限制对边缘设备(如DGX-Spark)提出了进一步挑战。为解决从云到边缘设备的这些多样化硬件瓶颈,我们提出了一个全栈推理流水线,将高效算法设计与优化的算子实现相结合。在算法上,我们引入了一个跨分辨率的两阶段生成调度器,利用扩散的逐步特性:早期低分辨率步骤快速建立全局布局,而后期高分辨率步骤专注于局部和感知细节的细化。这两个阶段通过一个学习的潜变量到潜变量映射模块连接,完全消除了跨不同分辨率进行分辨率转换时计算昂贵的VAE解码-重编码循环。在算子实现方面,我们部署了一个递归自我改进(RSI)循环,搜索内核融合和内存布局,并同时评估延迟和数值一致性。这些优化共同实现了高达30倍的端到端加速和20%的显存降低:在8xGB200节点上,一个5秒的1344x768带音频视频的生成速度比实时快3.5倍,并且在单台DGX Spark上不到一分钟即可完全驻留于显存中。
英文摘要
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.