SPEED:用于高质量视频帧插值的单步像素扩散
SPEED: One-Step Pixel Diffusion for High-quality Video Frame Interpolation
浏览论文内容
中文总结 AI 辅助
研究针对视频帧插值中现有扩散模型的问题,提出SPEED单步像素扩散框架。采用渐进多阶段架构、仅噪声更新注意力机制和漂移感知时间步采样策略,实现高质量单步推理,在多个基准测试中性能超越现有方法。
中文摘要 AI 辅助
尽管扩散模型在视频帧插值(VFI)中取得了成功,但现有方法仍存在两个关键限制。一是潜在扩散在从潜在表示重建图像回像素空间时不可避免地会丢失细粒度细节;二是多步采样会导致过高的内存消耗和推理延迟。为解决这些问题,我们提出了SPEED,一种用于高质量VFI的单步像素扩散框架。具体而言,SPEED采用具有动态补丁缩放的渐进多阶段架构来有效学习多尺度运动、结构和外观表示。此外,我们提出了一种新颖的仅噪声更新注意力机制,在减少近50%计算开销的同时防止干净条件帧的语义退化。我们还引入了一种漂移感知时间步采样策略和定制训练目标,以直接在像素空间中预测图像,实现单步推理且不影响生成帧的质量。大量实验表明,SPEED实现了领先性能。在SNU - FILM上,SPEED将LPIPS降低了8.8%,推理速度提高了63.3%,内存使用降低了10.6%。在具有挑战性的4K基准测试中,其LPIPS比先前方法高出多达51.5%。
英文摘要
Despite the success of diffusion models in Video Frame Interpolation (VFI), existing methods still suffer from two critical limitations. First, latent diffusion inevitably loses fine-grained details when reconstructing images from latent representations back to the pixel space. Second, multi-step sampling incurs prohibitive memory consumption and inference latency. To address these issues, we propose SPEED, a one-step pixel diffusion framework for high-quality VFI. Specifically, SPEED employs a progressive multi-stage architecture with dynamic patch scaling to effectively learn multi-scale motion, structural, and appearance representations. Furthermore, we propose a novel Noise-Update-Only Attention mechanism to prevent semantic degradation of the clean condition frames while reducing the computational overhead by nearly 50%. Besides, we introduce a Drift-aware Timestep Sampling strategy coupled with a tailored training objective to directly predict images in the pixel space, enabling one-step inference without compromising the quality of the generated frames. Extensive experiments show that SPEED achieves state-of-the-art performance. On SNU-FILM, SPEED reduces LPIPS by 8.8% while delivering 63.3% faster inference and 10.6% lower memory usage. On challenging 4K benchmarks, it further surpasses prior methods by up to 51.5% in LPIPS.