基于视觉-语义事件链条件的物理可信视频生成
Physically Plausible Video Generation via Visual-Semantic Chain-of-Events Conditioning
浏览论文内容
中文总结 AI 辅助
本文提出视觉-语义事件链条件框架,将物理可信视频生成重新表述为以事件为中心的生成,经多数据集实验验证其生成视频的物理可信性更优。
中文摘要 AI 辅助
物理可信视频生成(PPVG)旨在合成符合物理原理的视频,但由于自然语言条件的规定不足,该任务仍具挑战性。先进的思维链(CoT)框架会用物理知识增强提示,但这类提示整体描述物理现象,忽略了中间状态和过渡动态。本文将PPVG重新表述为以事件为中心的生成,将物理演化表示为因果关联且受物理约束的事件链。我们的框架包含三个关键模块:(1)物理驱动的事件链推理,该模块将物理现象分解为因果关联的事件,由演化的场景图表示,公式推导的物理量被绑定到相关对象和交互,表征每个事件过渡的方向和大小;(2)过渡感知的路由关键帧条件,该模块将每个事件路由到专门的关键帧合成算子,用于外观变化或对象变换,连续关键帧在去噪过程中作为残差引导注入,实现事件边界状态间的平滑视觉过渡;(3)物理注入的对比语义引导,该模块为无分类器引导构建物理感知的正提示和反事实负提示,引导生成趋向可信动态,远离违反物理的对应结果。在PhyGenBench、VideoPhy、PhyWorldBench和Physics-IQ上的实验表明,我们的框架在不同领域生成的视频具有更优的物理可信性。
英文摘要
Physically Plausible Video Generation (PPVG) seeks to synthesize videos consistent with physical principles, yet remains challenging due to underspecified natural language conditioning. Advanced chain-of-thought (CoT) frameworks augment prompts with physical knowledge. However, such prompts describe physical phenomena holistically, overlooking intermediate states and transition dynamics. In this paper, we reformulate PPVG as event-centric generation by representing physical evolution as a chain of causally connected and physically constrained events. Our framework comprises three key modules: (1) Physics-driven Event Chain Reasoning. This module decomposes physical phenomena into causally connected events represented by evolving scene graphs. Formula-derived physical quantities are bound to relevant objects and interactions, characterizing the direction and magnitude of each event transition. (2) Transition-aware Routed Keyframe Conditioning. This module routes each event to a specialized keyframe synthesis operator for appearance variation or object transformation. Consecutive keyframes are injected as residual guidance during denoising, enabling smooth visual transitions between event-boundary states. (3) Physics-injected Contrastive Semantic Guidance. This module constructs physics-informed positive and counterfactual negative prompts for classifier-free guidance, steering generation toward plausible dynamics and away from physics-violating counterparts. Experiments on PhyGenBench, VideoPhy, PhyWorldBench, and Physics-IQ demonstrate that our framework generates videos with superior physical plausibility across diverse domains.
发表机构
- College of Electronics and Information Engineering, Sichuan University(四川大学电子信息学院)
- School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
- School of Computer Science, University of Adelaide(阿德莱德大学计算机学院)
- School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。