SNM-VFI:对称非线性运动引导的生成式视频帧插值
SNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame Interpolation
浏览论文内容
中文总结 AI 辅助
提出无需训练的SNM-VFI框架,结合预训练光流与视频扩散模型,用对称非线性运动模型引导生成过程,经多基准评估实现了优质视频帧插值效果。
中文摘要 AI 辅助
我们提出了SNM-VFI(Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation,对称非线性运动引导的生成式视频帧插值),这是一种无需训练的框架,用于结合预训练光流模型和视频扩散模型实现可控制运动的生成式视频帧插值。与传统基于扩散的VFI方法从随机噪声合成中间帧不同,SNM-VFI通过对称非线性运动模型生成的对应感知帧引导生成过程。具体而言,我们首先利用预训练光流模型构建多帧基于非线性流的中间帧和置信度图,随后将这些流引导帧编码为潜在先验,以初始化并迭代引导预训练视频扩散模型,使该模型在保留密集运动对应关系的同时提升感知真实性。为进一步提升输出质量,我们采用置信度图在遮挡区域、物体边界等不确定区域,将结构可靠的基于流的预测与扩散生成的细节进行融合。在DAVIS、Sintel和KITTI等具有挑战性的基准上开展的大量评估表明,SNM-VFI在不同运动场景下实现了出色的感知质量、具有竞争力的重建精度以及稳健的时间一致性。
英文摘要
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model. Specifically, we first utilize a pre-trained optical flow model to construct multi-frame nonlinear flow-based intermediate frames and confidence maps. These flow-guided frames are then encoded as latent priors to initialize and iteratively guide a pre-trained Video Diffusion model, enabling the diffusion model to preserve dense motion correspondence while improving perceptual realism. To further enhance output quality, we employ confidence maps to fuse structurally reliable flow-based predictions with diffusion-generated details in uncertain regions such as occlusions and object boundaries. Extensive evaluations on challenging benchmarks, including DAVIS, Sintel, and KITTI, demonstrate that SNM-VFI achieves strong perceptual quality, competitive reconstruction accuracy, and robust temporal coherence across diverse motion scenarios.
发表机构
- Qualcomm(高通公司)
机构由 AI 辅助整理,请以论文原文为准。