AI 中文总结
研究旨在弥合移动视频扩散质量差距,提出将5B参数视频扩散Transformer通过循环重公式化等方法高效部署在移动硬件上,经多种优化,成为首个可商用移动设备部署的该规模模型,取得新的端到端延迟和分数,建立移动视频生成新水平。
AI 中文摘要
视频扩散的最新进展源于将基于Transformer的架构扩展到数十亿参数,显著提高了视觉保真度和运动连贯性。相比之下,现有的移动视频扩散模型参数预算相对较小,限制了生成质量。本文表明高质量移动视频生成无需小模型。通过循环重公式化和结构化压缩,可在内存受限的移动硬件上高效部署5B参数的服务器级视频扩散Transformer。从Wan2.2 - 5B开始,利用循环蒸馏框架,结合因果线性注意力,模型推理时像RNN一样运行并保持时间连贯性。还提出基于二进制头门的可学习注意力头修剪方法,结合采样步长蒸馏和内存优化的VAE解码,MobileWan成为首个可在商用移动设备上部署的5B规模视频扩散模型,实现了新的端到端延迟和VBench分数,在移动视频生成方面建立了新的技术水平。
英文摘要
Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coherence. In contrast, existing mobile video diffusion models remain limited to relatively small parameter budgets, typically 0.4-1.8B, restricting generation quality. In this work, we show that high-quality mobile video generation does not require small models. Instead, we demonstrate that a server-scale 5B-parameter video diffusion transformer can be deployed efficiently on memory-constrained mobile hardware through recurrent reformulation and structured compression. Starting from Wan2.2-5B, we rely on a recurrence distillation framework that converts video generation into a chunk-wise autoregressive process with constant-memory attention computation. Combined with causal linear attention, the model operates as an RNN at inference time while preserving temporal coherence across chunks. We further propose a learnable attention head pruning method based on binary per-head gates optimized end-to-end using a noise-biased sparsity objective and distillation-based finetuning. Together with sampling-step distillation and memory-optimized VAE decoding, MobileWan becomes the first 5B-scale video diffusion model deployable on a commercial mobile device. Our system generates 5-second 480x832 videos at 16 FPS in 20 seconds end-to-end latency, achieving a VBench score of 83.79 and establishing a new state of the art in mobile video generation. Please find the released DiT checkpoint and the sampling code in the project page: https://qualcomm-ai-research.github.io/MobileWan