BeyondMasks:评估视频对象移除中的因果与物理一致性
BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
浏览论文内容
中文总结 AI 辅助
该研究推出BeyondMasks基准及CORE评估协议,针对现有视频对象移除评估仅关注局部掩码保真度的问题,填补了视觉合理性与因果正确性的差距。
中文摘要 AI 辅助
生成式视频模型的最新进展显著提升了视频对象移除的视觉真实感,但评估协议仍聚焦于掩码区域的保真度,将移除视为局部修复。在真实场景中,对象移除是一种因果干预:消除对象还需移除其诱导的物理效应,如阴影、反射、光照变化、半透明效果及动态痕迹。现有基准缺乏对齐的干净参考,或仅局限于简化的合成设置,阻碍了对因果一致性的系统评估。我们推出BeyondMasks,这是一个用于因果一致视频对象移除的配对基准,包含时间对齐的合成与真实世界视频对及干净背景参考。该数据集涵盖多样的光度、几何、体积和动态交互,支持基于掩码和指令驱动的编辑。我们进一步提出CORE,一种基于结构化视觉语言模型的评估协议,联合测量对象消失与后效一致性,比现有指标更贴合人类判断。对当前最优方法的基准测试显示,尽管掩码区域保真度高,但在移除次级物理效应时存在系统性失败,暴露了视觉合理性与因果正确性之间的差距。BeyondMasks将视频对象移除重新定义为因果场景一致性而非局部重建,并为其评估提供了统一框架。
英文摘要
Recent advances in generative video models have significantly improved visual realism in video object removal, yet evaluation protocols still focus on masked region fidelity, treating removal as local inpainting. In real scenes, object removal is a causal intervention: eliminating an object also requires removing its induced physical effects, such as shadows, reflections, illumination changes, translucency, and dynamic traces. Existing benchmarks lack aligned clean references or remain limited to simplified synthetic settings, preventing systematic evaluation of causal consistency. We introduce BeyondMasks, a paired benchmark for causally consistent video object removal, consisting of temporally aligned synthetic and real world video pairs with clean background references. The dataset spans diverse photometric, geometric, volumetric, and dynamic interactions, and supports both mask based and instruction driven editing. We further propose CORE, a structured vision language model based evaluation protocol that jointly measures object disappearance and after effect consistency, aligning more closely with human judgments than existing metrics. Benchmarking state of the art methods reveals systematic failures in removing secondary physical effects despite high masked region fidelity, exposing a gap between visual plausibility and causal correctness. BeyondMasks reframes video object removal as causal scene consistency rather than local reconstruction and provides a unified framework for its evaluation.
发表机构
- Bilkent University(毕尔肯大学)
- Koç University(科克大学)
- Hacettepe University(哈杰泰佩大学)
- KUIS AI Center(KUIS人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。