arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视频帧插值中用于序列建模的跟随运动方法

Following Motion for Sequential Modeling in Video Frame Interpolation

Jaehyun Park, Nam Ik Cho

arXiv 2608.22861首次发表:更新:

发表机构

IPAI, Seoul National University(首尔国立大学 IPAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对SSMs预定义扫描顺序难以建模VFI动态运动轨迹的问题,提出MGMVFI模型,通过运动引导序列化与上下文合成等组件实现SOTA性能,为视频插值序列建模开辟新方向。

AI 中文摘要

状态空间模型(SSMs)已成为视频帧插值(VFI)领域极具潜力的架构,因其能以线性计算复杂度捕获长程依赖关系。然而,其预定义的扫描顺序限制了对VFI问题中固有动态运动轨迹的建模效果。为解决这一挑战,我们提出用于视频帧插值的运动引导Mamba(MGMVFI),这是一种专门针对VFI定制的选择性状态空间模型变体。MGMVFI引入运动引导序列化(MGS),利用光流为SSM定义运动自适应的一维输入顺序,使因果状态更新与语义相关的token对齐,实现运动一致的特征传播,尤其适用于大尺度动态运动。此外,为缓解光流估计不准确导致的不可靠特征表示,我们引入上下文合成,利用周围空间上下文实现鲁棒的帧间特征合成。这些组件无缝集成在我们定制的Mamba架构中,该架构还采用轻量细化模块以更低计算成本增强局部细节重建。在标准VFI基准上的大量实验表明,MGMVFI实现了SOTA性能,尤其在复杂动态运动场景中表现优异,从而为视频插值中的序列建模开辟了新方向。

英文摘要

State Space Models (SSMs) have surfaced as a promising architecture in Video Frame Interpolation (VFI), as they can capture long-range dependencies with linear computational complexity. However, their predefined scanning order limits their effectiveness in modeling the dynamic motion trajectories inherent in VFI problems. To tackle this challenge, we propose Motion-Guided Mamba for Video Frame Interpolation (MGMVFI), an adaptation of the selective state space model tailored explicitly for VFI. MGMVFI introduces Motion-Guided Serialization (MGS), which leverages optical flow to define a motion-adaptive 1D input order for the SSM. This aligns the causal state updates with semantically related tokens, enabling motion-consistent feature propagation, particularly for large and dynamic motions. Additionally, to mitigate the unreliable feature representations caused by inaccurate optical flow estimates, we introduce contextual synthesis that utilizes the surrounding spatial context for robust inter-frame feature synthesis. These components are seamlessly integrated within our tailored Mamba architecture, which also employs a lightweight refinement block to enhance local detail reconstruction at a reduced computational cost. Extensive experiments on standard VFI benchmarks demonstrate that MGMVFI achievesstate-of-the-artperformance,particularly on complex and dynamic motions, thereby establishing a new direction for sequence modeling in video interpolation.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑