arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31247cs.CVcs.AIcs.CRcs.MM

多视角图像集中的几何不一致性定位

Geometric Inconsistency Localization in Multi-View Image Sets

Xander Staelens, Albéric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert, Glenn Van Wallendael

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出DeformView数据集和DEFECt3R分类器,用于在宽基线多视角图像中定位几何不一致性,显著提升取证任务性能并减少误报。

中文摘要 AI 辅助

新视角合成(NVS)模型可以从不同视角生成同一场景的真实感新视图。然而,这些生成的视图之间并不总是几何一致的。多视角(MV)一致性已显示出作为评估这些NVS模型工具的潜力。然而,其在多媒体取证中的潜力在很大程度上仍未得到探索,特别是在宽基线图像对中定位几何不一致性方面。为了推动这一方向的研究,我们引入了DeformView,一个具有几何不一致性像素级标注的宽基线MV数据集。利用DeformView,我们评估了最先进的MV一致性评分方法,并表明为NVS评估开发的方法在几何不一致性定位的取证任务中迁移效果不佳。为解决这一局限,我们提出了DEFECt3R,一种轻量级基于学习的分类器,利用跨视角特征关系在像素级别定位几何不一致性。通过从显式监督中学习,包括来自几何一致但变形视图的难负样本,DEFECt3R相比现有的一致性评分方法提高了定位性能,并大幅减少了误报。消融实验进一步表明,特征表示和对应质量都对定位性能有所贡献。总体而言,我们的研究结果表明,MV几何一致性是多媒体取证中一个有前景但尚未充分探索的信号,并为宽基线MV图像对中的几何不一致性定位建立了基准和基线。代码和数据集可在以下https URL获取。

英文摘要

Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image pairs. To enable research in this direction, we introduce DeformView, a wide-baseline MV dataset with pixel-level annotations of geometric inconsistencies. Using DeformView, we evaluate state-of-the-art MV consistency-scoring methods and show that approaches developed for NVS evaluation transfer poorly to the forensic task of geometric inconsistency localization. To address this limitation, we propose DEFECt3R, a lightweight learning-based classifier that uses cross-view feature relationships to localize geometric inconsistencies at the pixel level. By learning from explicit supervision, including hard negatives from geometrically consistent yet deformed views, DEFECt3R improves localization performance and substantially reduces false positives compared to existing consistency-scoring methods. Ablation experiments further show that both feature representations and correspondence quality contribute to localization performance. Overall, our findings demonstrate that MV geometric consistency is a promising yet underexplored signal for multimedia forensics and establish a benchmark and baseline for geometric inconsistency localization in wide-baseline MV image pairs. Code and dataset are available at https://github.com/IDLabMedia/DeformView-DEFECt3R

发表机构

  • IDLab, Ghent University - imec(根特大学-imec IDLab)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑