arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sol-H3:面向云与边缘的Sol-Engine上MiniMax-H3推理加速的递归自我改进

Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge

Yitong Li, Jincheng Yu, Junsong Chen, Haopeng Li, Shuchen Xue, Haozhe Liu, Ping Luo, Song Han, Enze Xie

arXiv 2609.35110首次发表:更新:

发表机构

NVIDIA Research, Efficient AI Team & Singapore Lab(NVIDIA Research,高效AI团队与新加坡实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对MiniMax-H3视频扩散模型在云与边缘设备的推理瓶颈,提出全栈优化流水线,含跨分辨率调度与递归自我改进内核搜索,实现30倍加速和20%显存降低。

AI 中文摘要

视频扩散模型正在快速扩展并展现出更强的生成能力。在这些最新进展中,MiniMax-H3作为一个高能力、生产级的开源模型脱颖而出。然而,其330亿参数和逐步迭代去噪过程引入了巨大的计算开销。因此,其实际生产受到云部署(如NVIDIA-GB200)中生成延迟的阻碍,同时严格的显存限制对边缘设备(如DGX-Spark)提出了进一步挑战。为解决从云到边缘设备的这些多样化硬件瓶颈,我们提出了一个全栈推理流水线,将高效算法设计与优化的算子实现相结合。在算法上,我们引入了一个跨分辨率的两阶段生成调度器,利用扩散的逐步特性:早期低分辨率步骤快速建立全局布局,而后期高分辨率步骤专注于局部和感知细节的细化。这两个阶段通过一个学习的潜变量到潜变量映射模块连接,完全消除了跨不同分辨率进行分辨率转换时计算昂贵的VAE解码-重编码循环。在算子实现方面,我们部署了一个递归自我改进(RSI)循环,搜索内核融合和内存布局,并同时评估延迟和数值一致性。这些优化共同实现了高达30倍的端到端加速和20%的显存降低:在8xGB200节点上,一个5秒的1344x768带音频视频的生成速度比实时快3.5倍,并且在单台DGX Spark上不到一分钟即可完全驻留于显存中。

英文摘要

Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑