arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10479cs.CV

连接事件流与DiT:事件引导的视频帧插值

Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

  • The University of Tokyo(东京大学)
  • South China University of Technology(华南理工大学)
  • Wuhan University(武汉大学)
  • Singapore Management University(新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

Guixu Lin, Yuyang Yu, Xiang Ji, Linyao Chen, Zhengwei Yin, Mengshun Hu, Mingdeng Cao, Shengfeng He, Yinqiang Zheng

AI总结:

本文提出基于适配器的框架,将事件衍生提示融入预训练图像转视频扩散模型,利用IWEs与双向稀疏光流引导生成,提升视频帧插值质量,效果优于现有SOTA方法。

AI中文摘要:

潜在扩散模型近期通过合成输入图像间的中间帧推进了视频帧插值技术,但处理大时间间隔与复杂运动仍具挑战,常导致运动模糊、结构扭曲及时间不一致问题。事件相机提供的高时间分辨率运动提示,适合填补这些间隔并提升插值质量。为利用该优势且无需从头训练事件辅助模型,本文提出一种基于适配器(adapter)的框架,以最小架构改动将事件衍生提示融入预训练的图像转视频扩散模型。具体而言,该方法利用图像扭曲事件(IWEs)与双向稀疏光流,在生成过程中提供时空对齐的引导;通过将这些事件引导的结构与运动提示注入扩散过程,减少插值伪影并提升重建保真度与时间一致性。在真实及合成基准上的实验结果表明,本文方法始终优于现有最先进方法,项目页面位于该https URL。

英文摘要:

Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling large temporal gaps and complex motion remains challenging, often resulting in motion blur, structural distortions, and temporal inconsistencies. Event cameras provide high-temporal-resolution motion cues that are well suited for bridging these gaps and improving interpolation quality. To exploit this advantage without training an event-assisted model from scratch, we propose an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes. Specifically, our method leverages Image Warped Events (IWEs) and bidirectional sparse optical flow to provide spatially and temporally aligned guidance during generation. By injecting these event-guided structural and motion cues into the diffusion process, our approach reduces interpolation artifacts and improves both reconstruction fidelity and temporal coherence. Experimental results on real and synthetic benchmarks show that our method consistently outperforms existing state-of-the-art approaches. The project page is at https://joseph-lin-tech.github.io/BridgeEventDiT-VFI/.

补充信息

↑