arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07120cs.CV

匹配歧义下视频帧插值的多假设流估计

Multiple Hypothesis Flow Estimation for Video Frame Interpolation under Matching Ambiguity

Zibo Su, Jing Kong, Ruixing Wang, Zhanhe Zhang, Kun Wei

首次发表
浏览论文内容

中文总结 AI 辅助

针对视频帧插值中匹配歧义导致的重影等问题,提出多假设流估计框架,保留前K个候选对应并选优,在多基准上获最优LPIPS与DISTS

中文摘要 AI 辅助

许多基于流的视频帧插值(VFI)方法通过估计光流场、对两输入帧进行变形并融合变形后的观测结果来合成中间帧。这些潜在流场通常通过图像级重建监督学习得到,无直接流标注。在包含重复或随机纹理、旋转对称结构或带模糊的快速运动的歧义区域中,单个查询的匹配证据可能包含多个可比较且空间分离的峰值。尽管真实中间帧提供间接监督,但可能无法唯一确定歧义区域中的潜在对应关系。若多个位置提供多个合理匹配,单一流估计器只能保留一个位移并丢弃其余候选;若所选匹配不正确或与相邻像素的匹配不一致,变形会从不匹配位置采样内容,产生重影、结构失真或模糊。为解决此局限,我们提出多假设流估计框架,保留前K个候选对应关系并通过可靠性引导路由器为每个位置选择一个;每个假设从粗匹配锚点初始化,并通过以锚点为中心的局部注意力分别优化。帧合成因此以所选的流-外观假设为条件,而非候选的软组合。在提出的MA-HD基准及公开VFI基准上的实验表明,我们的方法在对比方法中取得了最佳的LPIPS和DISTS指标。

英文摘要

Many flow-based video frame interpolation (VFI) methods synthesize an intermediate frame by estimating optical flow fields, warping the two input frames, and blending the warped observations. These latent flow fields are typically learned through image-level reconstruction supervision without direct flow annotations. In ambiguous regions containing repetitive or stochastic textures, rotating symmetric structures, or fast motion with blur, the matching evidence for a single query may contain multiple comparable and spatially separated peaks. Although the ground-truth intermediate frame provides indirect supervision, it may not uniquely identify the latent correspondence in ambiguous regions.When several locations provide multiple plausible matches, a single-flow estimator can retain only one displacement and discard the remaining candidates. If the selected match is incorrect or inconsistent with those of neighboring pixels, warping samples content from mismatched locations, producing ghosting, structural distortion, or blur.To address this limitation, we propose a multiple hypothesis flow estimation framework that preserves top-K candidate correspondences and selects one per location through a reliability-guided router. Each hypothesis is initialized from a coarse matching anchor and refined separately through anchor-centered local attention. Frame synthesis is thus conditioned on one selected flow-appearance hypothesis rather than a soft combination of candidate motions.Experiments on the proposed MA-HD benchmark and public VFI benchmarks show that our method achieves the best LPIPS and DISTS among the compared methods.

↑