UnStep:无需蒸馏训练即可用更少步数加速因果视频扩散
UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation
AI总结:
提出无需训练的UnStep包装器,通过减少DiT步数、限制KV缓存窗口及截断SVD等推理机制,显著加速因果视频扩散,在H100和GB200上分别达到50和77 FPS且无质量损失。
AI中文摘要:
将双向多步视频扩散变换器蒸馏为少步因果模型已成为流式视频生成的常见方法。虽然这些少步学生模型比其蒸馏来源的教师模型快得多,但对于实时生成而言仍然较慢。在这项工作中,我们提出了UnStep,一种无需训练的包装器,通过在推理时以比蒸馏时更少的扩散变换器(DiT)步数运行少步因果视频模型,并限制注意力KV缓存中保留的时间窗口,从而加速推理。我们提出了两种仅推理的机制来恢复因步数减少和注意力窗口化而损失的质量:将生成的潜在帧重新加噪至接近干净的水平,并重用现有的干净缓存通道来细化它们;以及对DiT注意力值和输出投影应用截断SVD。我们还通过为DiT和VAE解码器提供保持质量的运行时栈来加速推理,包括更高效的注意力调用和KV索引、使用缓存系数的融合Triton RoPE,以及采用优化内存布局、精度和卷积内核的VAE解码。通过减少计算和优化运行时栈,UnStep为因果视频扩散设定了新的吞吐量水平,其运行速度远超当前方法,在单个H100上达到50 FPS且无质量损失,在GB200上达到77 FPS,且无需重新训练。
英文摘要:
Distilling bidirectional multi-step video diffusion transformers into few-step causal models has become a common approach for streaming video generation. While these few-step students are significantly faster than the teachers they are distilled from, they remain slow for real-time generation. In this work we present UnStep, a training-free wrapper that accelerates few-step causal video models at inference by running them with fewer diffusion transformer (DiT) steps than during distillation and limiting the temporal window retained in the attention KV cache. We propose two inference-only mechanisms to recover quality lost by step reduction and attention windowing: renoising the generated latent frames to a near-clean level and reusing the existing clean-cache pass to refine them, and applying truncated SVD to the DiT attention value and output projections. We also accelerate inference with a quality-preserving runtime stack for the DiT and VAE decoder, including more efficient attention calls and KV indexing, fused Triton RoPE with cached coefficients, and VAE decoding with optimized memory layout, precision, and convolution kernels. By reducing computation and optimizing the runtime stack, UnStep sets a new throughput regime for causal video diffusion, by running substantially faster than current methods, reaching 50 FPS on a single H100 without quality loss, and 77 FPS on GB200, all without retraining.