VOR-Bench:一种人类感知驱动的视频对象移除基准
VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal
- Beijing University of Posts and Telecommunications(北京邮电大学)
- China Telecom Artificial Intelligence Technology Co. Ltd(中国电信人工智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视频对象移除评估中参考不可靠和指标与人类偏好错位的问题,提出VOR-Bench基准,包含VORD数据集、rMPAF采集框架和VOR-MDSM评分模型,实现与人类感知高度一致(ρ>0.9)的评估。
AI中文摘要:
尽管视频对象移除(VOR)在其中扮演着关键角色,现有的评估范式仍面临两个关键限制:参考的可疑性以及传统指标与人类偏好之间的错位。为应对这些挑战,我们引入了VOR-Bench,通过三个集成组件推进VOR评估。首先,我们提出了VOR数据集(VORD),这是首个同时提供成对编辑视频和涂鸦掩码的基准数据集。其独特优势在于多样化的数据谱系,涵盖模型生成、工具渲染和相机捕获的数据,确保在真实世界场景中的稳健评估。其次,我们开发了rMPAF,一种逼真的、支持运动的成对视频采集框架。通过结合基于图像的对象移除和微调视频生成模型的优势,rMPAF自动生成逼真的、运动连贯的成对视频。最后,我们提出了三个评估维度,并引入了VOR-MDSM,这是首个专门为掩码引导的VOR设计的基于感知的VLM评分模型。它通过覆盖基本视觉属性和匹配细微的人类判断,弥合了算术指标与人类感知之间的差距。大量实验表明,VOR-Bench产生的评估结果与人类感知高度一致,与主观评估实现了显著的相关性(ρ > 0.9)。我们将发布VOR-Bench及其文档,以确保完全可复现性。
英文摘要:
Despite its crucial role in video object removal (VOR), existing evaluation paradigms face two critical limitations: questionable references and a misalignment between tradi- tional metrics and human preference. To address these challenges, we introduce VOR- Bench, which advances VOR evaluation through three integrated components. First, we present the VOR Dataset (VORD), the first benchmark dataset providing both paired edited videos and graffiti masks. Its unique strength lies in a diverse data spectrum, which encompasses model-generated, tool-rendered, and camera-captured data, ensuring robust assessment across real-world scenarios. Second, we develop rMPAF, a realistic Motion- capable Paired-video Acquisition Framework. By combining the strengths of image- based object removal and fine-tuned video generation models, rMPAF automatically generates realistic, motion-coherent paired videos. Finally, we propose three evaluation dimensions and introduce VOR-MDSM, the first perception-driven VLM-based scoring model specifically designed for mask-guided VOR. It bridges the gap between arithmetic metrics and human perception by covering the essential visual attributes and matching nuanced human judgment. Extensive experiments demonstrate that VOR-Bench yields evaluation results that align closely with human perception, achieving a remarkable cor- relation (\r{ho} > 0.9) with subjective assessments. We will release VOR-Bench along with its documentation to ensure full reproducibility.