发表机构
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频编辑中物理后果未充分探索的问题,提出物理反事实视频编辑(PCVE)任务,并引入VideoPhysEdit免训练流水线,通过物理场景重建显式推理,结合新基准PCVE-RigidBench和物理编辑分数,在刚体场景中实现高物理编辑准确性。
AI 中文摘要
近年来,视频编辑技术取得了显著进展,方法日益考虑编辑的视觉后果,如阴影和遮挡的变化。然而,编辑的物理后果,包括后续运动和交互的变化,仍较少被探索。我们将此问题定义为物理反事实视频编辑(PCVE),其目标是在给定源视频、物理编辑及其执行帧的情况下,生成描绘由此产生的运动和交互的反事实视频。PCVE具有挑战性,因为它需要理解场景物理并推断由物理干预引起的下游运动和交互,同时缺乏成对的事实和反事实数据以及专门的评估指标。我们引入了VideoPhysEdit,一种用于刚体场景中PCVE的新型免训练流水线。它通过一种新颖的物理场景重建方法使物理推理明确化,该方法恢复一个在模拟下重现观察到的运动和交互的场景,从而使流水线能够将物理编辑作为干预应用,并利用由此产生的轨迹来指导反事实视频生成。我们进一步构建了PCVE-RigidBench,一个带有成对源视频和反事实目标视频以及物理真实值的合成基准,并引入了物理编辑分数。VideoPhysEdit在物理编辑准确性上显著高于开源方法和商业模型,同时保持有竞争力的视觉保真度。其物理编辑分数为0.376,是所比较方法中唯一的正分数。对真实视频的定性比较进一步表明,VideoPhysEdit适用于现实世界场景,并且比所比较的方法更好地描绘了由编辑引起的下游运动和交互。代码:此https URL
英文摘要
Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored. We formulate this problem as physical counterfactual video editing (PCVE), which aims to generate a counterfactual video depicting the resulting motion and interactions given a source video, a physical edit, and its execution frame. PCVE is challenging because it requires understanding scene physics and inferring the downstream motion and interactions induced by a physical intervention, while paired factual and counterfactual data and dedicated evaluation metrics are lacking. We introduce VideoPhysEdit, a new training-free pipeline for PCVE in rigid-body scenes. It makes physical reasoning explicit through a novel physical scene reconstruction method that recovers a scene reproducing the observed motion and interactions under simulation, enabling the pipeline to apply physical edits as interventions and use the resulting trajectories to guide counterfactual video generation. We further construct PCVE-RigidBench, a synthetic benchmark with paired source and counterfactual target videos and physical ground truth, and introduce the Physical Edit Score. VideoPhysEdit achieves substantially higher physical edit accuracy than open-source methods and commercial models while maintaining competitive visual fidelity. Its Physical Edit Score is 0.376, the only positive score among the compared methods. Qualitative comparisons on real videos further show that VideoPhysEdit applies to real-world scenes and better depicts the downstream motion and interactions induced by the edits than the compared methods. Code: https://github.com/Hammour-steak/VideoPhysEdit