发表机构
School of Computer Science, Nanjing University; Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology; SenseTime; Shenzhen Technology University(南京大学计算机科学学院; 深圳理工大学人工智能研究院; 商汤科技; 深圳技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对固气动力学视频生成的相位耦合挑战,提出双分支架构的PAVG模型,构建含70万条轨迹的模拟语料库,生成的视频在运动贴合度等指标上优于现有方法。
AI 中文摘要
生成具有物理合理性的固气动力学视频颇具挑战,因为不同相位展现出截然不同的动力学特性,却又通过物理交互相互耦合。我们提出PAVG,一种面向固气动力学与交互的相位感知视频生成器。它采用双分支架构显式建模固体与气体的不同动力学,同时通过时空交叉注意力捕捉二者的物理交互。该设计使PAVG能够保留特定相位的运动特征,同时在各相位间生成物理一致的响应。为推进该任务,我们进一步构建了一个模拟语料库,涵盖超过70万条物理轨迹,涉及多样的固体、气体及固气交互场景。大量评估表明,与现有方法相比,我们的PAVG生成的视频在运动贴合度、物理合理性及视觉质量上均有提升。
英文摘要
Generating physically plausible videos for solid-gas dynamics is challenging as different phases exhibit distinct dynamics yet remain coupled through physical interactions. We present PAVG, a Phase-Aware Video Generator for solid-gas dynamics and interactions. It employs a dual-branch architecture to explicitly model the distinct dynamics of solids and gases, while spatiotemporal cross-attention captures their physical interactions. This design enables PAVG to preserve phasespecific motion characteristics while producing physically consistent responses across phases. To facilitate this task, we further construct a simulation corpus comprising over 700K physical trajectories across diverse solid, gas, and solid-gas interaction scenarios. Extensive evaluations demonstrate that our PAVG produces videos with improved motion adherence, physical plausibility, and visual quality compared with existing approaches.
Comments31 pages, 5 figures, 9 tables, including appendix