arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MobileWan:弥合移动视频扩散的质量差距

MobileWan: Closing the Quality Gap for Mobile Video Diffusion

Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv, Ioannis Lelekas, Noor Fathima, Spyridon Stasis, Hanno Ackermann, Boris van Breugel, Markus Nagel, Fatih Porikli, Animesh Karnewar, Amirhossein Habibian

arXiv 2607.06173首次发表:更新:

AI 中文总结

研究旨在弥合移动视频扩散质量差距,提出将5B参数视频扩散Transformer通过循环重公式化等方法高效部署在移动硬件上,经多种优化,成为首个可商用移动设备部署的该规模模型,取得新的端到端延迟和分数,建立移动视频生成新水平。

AI 中文摘要

视频扩散的最新进展源于将基于Transformer的架构扩展到数十亿参数,显著提高了视觉保真度和运动连贯性。相比之下,现有的移动视频扩散模型参数预算相对较小,限制了生成质量。本文表明高质量移动视频生成无需小模型。通过循环重公式化和结构化压缩,可在内存受限的移动硬件上高效部署5B参数的服务器级视频扩散Transformer。从Wan2.2 - 5B开始,利用循环蒸馏框架,结合因果线性注意力,模型推理时像RNN一样运行并保持时间连贯性。还提出基于二进制头门的可学习注意力头修剪方法,结合采样步长蒸馏和内存优化的VAE解码,MobileWan成为首个可在商用移动设备上部署的5B规模视频扩散模型,实现了新的端到端延迟和VBench分数,在移动视频生成方面建立了新的技术水平。

英文摘要

Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coherence. In contrast, existing mobile video diffusion models remain limited to relatively small parameter budgets, typically 0.4-1.8B, restricting generation quality. In this work, we show that high-quality mobile video generation does not require small models. Instead, we demonstrate that a server-scale 5B-parameter video diffusion transformer can be deployed efficiently on memory-constrained mobile hardware through recurrent reformulation and structured compression. Starting from Wan2.2-5B, we rely on a recurrence distillation framework that converts video generation into a chunk-wise autoregressive process with constant-memory attention computation. Combined with causal linear attention, the model operates as an RNN at inference time while preserving temporal coherence across chunks. We further propose a learnable attention head pruning method based on binary per-head gates optimized end-to-end using a noise-biased sparsity objective and distillation-based finetuning. Together with sampling-step distillation and memory-optimized VAE decoding, MobileWan becomes the first 5B-scale video diffusion model deployable on a commercial mobile device. Our system generates 5-second 480x832 videos at 16 FPS in 20 seconds end-to-end latency, achieving a VBench score of 83.79 and establishing a new state of the art in mobile video generation. Please find the released DiT checkpoint and the sampling code in the project page: https://qualcomm-ai-research.github.io/MobileWan

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑