arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17695cs.CV

基于流匹配模型的快速视频生成的幅度-方向解耦方法

Magnitude-Direction Decoupling for Fast Video Generation with Flow Matching Models

  • Nanjing University of Science and Technology(南京理工大学)
  • Huawei(华为)

机构由 AI 辅助整理,请以论文原文为准。

Haonan Xu, Feiyang Chen, Songkui Chen, Hongpeng Pan, Zhefeng Wang, Xinyu Duan, Baoxing Huai, Yang Yang

中文总结 AI 辅助

针对流匹配视频生成模型计算开销高的问题,提出MDD方法,通过解耦幅度与方向复用轻量模型,在Wan2.1上实现最高2.95倍加速,性能优于现有方法且保真度高。

中文摘要 AI 辅助

用于视频生成的流匹配模型实现了出色的性能,但由于迭代去噪,存在计算开销高的问题。实际上,原始模型并非在所有去噪步骤都必要,允许部分步骤使用轻量替代模型以加快采样速度。然而,直接使用缓存或轻量模型会偏离原始去噪轨迹,导致性能不佳。通过实证分析,我们发现轻量模型可稳健捕获原始模型输出的幅度分量,而缓存可提供可靠的方向引导。基于此见解,我们提出幅度-方向解耦(Magnitude-Direction Decoupling,MDD)方法,该方法自适应采用经方向校准的轻量模型作为原始模型的替代,以加速推理并有效修正去噪轨迹的偏差。此外,MDD通过在无分类器引导(classifier-free guidance,CFG)下复用幅度信息进一步降低推理成本。因此,MDD提供了一种更可靠、更轻量的采样加速解决方案。实验表明,MDD的性能优于现有加速方法,实现了可观的加速比(例如在Wan2.1上最高达2.95倍),同时保持了高视觉保真度和内容丰富度。

英文摘要

Flow matching models for video generation achieve impressive performance but suffer from high computational overhead due to iterative denoising. In fact, the original model is not necessary for all denoising steps, allowing some steps to use lightweight alternatives for faster sampling. However, directly using caching or lightweight models can deviate from the original denoising trajectory, resulting in suboptimal performance. Through empirical analysis, we find that lightweight models can robustly capture the magnitude components of the original model's output, while caching provides reliable directional guidance. Building on this insight, we propose the Magnitude-Direction Decoupling (MDD) method, which adaptively employs a direction-calibrated lightweight model as a substitute for the original model to accelerate inference and effectively correct deviations in the denoising trajectory. Moreover, MDD further reduces inference costs by reusing magnitude information under classifier-free guidance (CFG). As a result, MDD offers a more reliable and lightweight solution to accelerate sampling. Experiments show that MDD outperforms existing acceleration methods, delivering promising speedups (e.g., up to 2.95x on Wan2.1) while preserving high visual fidelity and content richness.

↑