基于智能体规划与图引导优化的物理感知视频生成
PhysPlan: Grounded Physical State Reasoning and Graph-Guided Optimization for Physically Plausible Video Generation
浏览论文内容
中文总结 AI 辅助
针对视频扩散模型缺乏物理理解的问题,提出PhysPlan框架,通过智能体规划和对象中心梯度路由优化,在多个基准上显著提升物理感知视频生成性能。
中文摘要 AI 辅助
视频扩散模型(VDM)在合成高保真、逼真的视频内容方面已展现出卓越能力。然而,它们从根本上缺乏对物理定律的内在理解,经常产生视觉上吸引人但因果逻辑上不合理的序列,其特征是结构幻觉和物理上不合理的动态。通过无需训练(免训练)的测试时优化来注入物理感知是一种有前景的替代方案,但现有方法依赖于全局梯度更新和刚性调度启发式方法,这些方法会无意中破坏被动背景,并且无法对复杂的动态状态变化进行建模。为解决这一问题,我们提出了PhysPlan,一种新颖的无需训练(免训练)的引导框架,它将范式从随机视觉插值转变为智能体物理模拟。首先,视觉语言模型(VLM)作为迭代认知模拟器,将多模态输入分解为视觉思维链(Chain-of-Visual-Thought),以创建运动学轨迹和3D深度几何的多模态表示。其次,这些信号驱动以对象为中心的测试时优化。与先前依赖全局梯度和刚性调度启发式方法的无需训练方法不同,PhysPlan引入了对象中心梯度路由(Object-Centric Gradient Routing)来隔离运动学修改并完全锁定被动环境。此外,我们的动力学强度分析(Kinetic Intensity Profiling)动态参数化框架超参数,以适应物理变形严重程度的变化。在PhyGenBench和Physics-IQ基准上的广泛评估表明,PhysPlan显著优于基础性和可控性视频扩散模型基线,为提升视频生成的物理理解提供了一种有前景的方法。
英文摘要
Video diffusion models (VDMs) synthesize photorealistic content, yet they often fail to follow the course that a physical phenomenon should take within a given scene. Recent training-free methods let a vision-language model (VLM) plan the phenomenon and guide a frozen VDM toward the plan; however, such plans are derived from the prompt and consumed as whole keyframes or trajectories, which leaves unspecified where the consequences land in the observed scene and turns incidental visual details into optimization targets. We observe that a phenomenon specified in words unfolds as sparse, local changes to the physical state of the observed scene. Building on this observation, we present PhysPlan, a training-free image-to-video framework that represents a phenomenon as a grounded state graph and uses this graph to decide what, where, and when the guidance constrains. Grounded Physical State Reasoning decomposes the phenomenon into physical deltas, each stating which objects change, to what state, and by which physical rule, and translates each delta into graph edits, verified by deterministic checks, that leave all other objects unchanged. Graph-Guided Test-Time Optimization renders a keyframe for each state, measures the denoised estimates only along the properties selected by the edits, and concentrates the update on the edited objects. On PhyGenBench and Physics-IQ, PhysPlan raises its base model from 0.52 to 0.77 and from 27.1 to 38.2, surpassing the strongest prior I2V method (0.60 and 34.6), and lowers FVD by over 20%. Project page: https://physplan.github.io
发表机构
- University of Science, Ho Chi Minh City(胡志明市理科大学)
- Vietnam National University, Ho Chi Minh City(越南国立大学胡志明市分校)
- Monash University(莫纳什大学)
- University of Dayton(代顿大学)
机构由 AI 辅助整理,请以论文原文为准。