arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysFlow:面向运动可控视频生成的物理感知光流

PhysFlow: Physics-Aware Optical Flow for Motion Controllable Video Generation

Cong Wang, Hanxin Zhu, Yonglin Tian, Jiayi Luo, Ruiqi Song, Boyi Sun, Long Chen, Zhibo Chen

arXiv 2609.08215首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Zhongguancun Academy; School of Information Science and Technology, University of Science and Technology of China; Beihang University(中国科学院自动化研究所; 中国科学院大学; 中关村 Academy; 中国科学技术大学信息科学技术学院; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhysFlow提出两阶段框架,先物理感知生成光流再据此合成外观,并构建含10K物体、50K序列的PhysVideo数据集,以提升生成视频的物理合理性与视觉保真度。

AI 中文摘要

视频生成模型因其生成视觉上引人注目的视频的能力而近来受到广泛关注,然而确保物理一致且合理的动态仍然是一个基本挑战,这推动了关于视频生成中物理真实性的日益增长的研究方向。为了解决这一挑战,基于物理规律主要编码在运动模式中的事实,我们提出了PhysFlow,一种新颖的两阶段框架,通过将视频生成分解为运动感知的光流生成以及随后的运动条件化外观合成,来提高生成视频的物理合理性。具体来说,PhysFlow由一个名为PA-Flow的物理感知光流视频生成器和一个名为FlowRender的光流引导视频生成器组成。在第一阶段,PA-Flow采用物理感知注意力模块来分别建模运动属性和材料属性如何影响全局运动和局部变形,并生成光流视频作为运动的显式表示。在第二阶段,FlowRender利用解耦的运动表示作为引导来合成真实的纹理和外观,最终产生物理上合理的视频。为了进一步支持带有显式物理监督的模型训练,我们构建了PhysVideo,一个基于物理引擎和3D-GS渲染生成的基于物理的视频数据集,包含10K个前景物体和50K个具有运动及材料属性标注的真实视频序列。大量实验表明,与现有方法相比,我们提出的PhysFlow生成的视频具有优越的物理合理性,同时保持了较高的视觉保真度。

英文摘要

Video generation models have recently attracted substantial attention for their ability to generate visually compelling videos, yet ensuring physically consistent and plausible dynamics still remains a fundamental challenge, driving a growing line of research on physical realism in video generation. To address this challenge, motivated by the fact that physical regularities are primarily encoded in motion patterns, we propose PhysFlow, a novel two-stage framework for improving the physical plausibility of generated videos by decomposing video generation into motion-aware optical flow generation followed by motion-conditioned appearance synthesis. Specifically, PhysFlow consists of a physics-aware optical-flow video generator called PA-Flow and a flow-guided video generator called FlowRender. During the first stage, PA-Flow employs a physics-aware attention module to model how motion attributes and material properties influence global motion and local deformation, respectively, and generates an optical flow video as an explicit representation of motion. In the second stage, FlowRender leverages the decoupled motion representation as guidance to synthesize realistic textures and appearances, ultimately producing the final physically plausible video. To further support model training with explicit physical supervision, we construct PhysVideo, a physics-based video dataset generated with a physics engine and 3D-GS rendering, containing 10K foreground objects and 50K realistic video sequences with annotations of motion and material properties. Extensive experiments demonstrate that our proposed PhysFlow generates videos with superior physical plausibility while maintaining high visual fidelity compared with existing methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑