EffectLearner:面向真实世界视频物体移除的世界感知物体-效应推理
EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
浏览论文内容
中文总结 AI 辅助
该研究提出EffectLearner框架,结合VLM物体-效应推理器与DiT视频擦除器,构建专属数据集EffectWorld并采用渐进式训练,在视频物体移除任务上实现更优性能。
中文摘要 AI 辅助
视频物体移除不仅需要消除目标物体,还需消除其诱导效应,同时保持高保真度和时空一致的修复效果。现有方法主要从预定义的效应类别和固定数据分布中隐式学习物体-效应对应关系,这限制了它们在复杂真实世界场景中的泛化能力,此类场景包含组合效应、空间分离或弱相关效应、长尾物理现象以及动态演化的交互作用。我们提出EffectLearner,这是一个经语义推理增强的框架,结合了基于VLM的物体-效应推理器与基于DiT的视频擦除器。在结构化效应分析提示的引导下,该推理器对经目标高亮处理的视频执行跨模态推理,提取紧凑的感知效应上下文,以此指导视频擦除器实现全面的物体-效应移除。运动感知掩码引导和运动一致性监督进一步提升了物体运动和场景动态演化下的移除覆盖范围和时空稳定性。为在具有挑战性的真实场景中充分利用该框架,我们还构建了EffectWorld,这是一个专为复杂物体诱导效应设计的配对视频数据集,并引入了渐进式训练课程,将常规监督与复杂效应数据相结合。在标准ROSE-Bench上,EffectLearner在大多数指标上优于现有基线方法,且在EffectWorld-Eval和更具挑战性的EffectWorld-Wild上均取得明显优势,证明其在复杂真实场景中能够提供高质量的视频物体移除效果。
英文摘要
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.