arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习用于视频生成的显式物理参数控制和基准测试

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Yanxun Li, Hao Wen, Bingze Song, Jiashu Zhu, Aiming Hao, Chubin Chen, Jintao Chen, Jiahong Wu, Xiangxiang Chu, Miao Wang

arXiv 2607.18924首次发表:更新:

发表机构

Beihang University; Alibaba Group(北京航空航天大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对图像到视频生成中物理参数控制问题,引入含密集物理参数化的数据集PhyParam-Dataset,提出物理引导的扩散模型PhyParam,建立基准PhyParam-Bench,通过实验证明该模型能在保持视觉保真度时提升物理一致性。

AI 中文摘要

图像到视频生成的最新进展提高了视觉真实感,使基于物理且可控的动力学成为未来世界模拟的重要一步。当前模型常生成合理运动,但不受明确物理原因可靠控制,实例级约束会泄漏或在多对象交互中纠缠。我们将此差距归因于两点缺失:大规模、细粒度物理参数化,以及正确将物理属性绑定到实例并强调动力学而非外观的模型设计。为弥合差距,我们引入PhyParam-Dataset,一个以交互为中心的包含130K个具有密集物理参数化的物理模拟视频的数据集。在此数据基础上,我们提出PhyParam,一种物理引导的图像到视频扩散模型,通过轻量级物理注意力路由机制以对象级力、质量、摩擦、恢复系数和场景级重力为条件,并通过语义结构特征空间监督进一步改进运动学习。我们还建立了PhyParam-Bench,一个用于图像到视频生成中物理定律一致性的基准,通过多层次协议评估时间动态、空间稳定性和语义-物理对齐。实验表明,PhyParam在保持高视觉保真度的同时提高了物理一致性,推进了图像到视频生成的显式刚体物理参数控制。我们将公开发布数据集、基准和代码以支持未来研究。

英文摘要

Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions. We attribute this gap to two missing pieces: large-scale, fine-grained physical parameterization, and model designs that correctly bind physical attributes to instances and emphasize dynamics over appearance. To bridge this gap, we introduce PhyParam-Dataset, an interaction-centric collection of 130K physically simulated videos with dense physical parameterization, including force vectors, object material properties, and environmental constants across five representative rigid-body motion types. Built on this data, we present PhyParam, a physics-guided image-to-video diffusion model that conditions on object-level forces, masses, friction, restitution, and scene-level gravity via a lightweight physical-attention routing mechanism, and further improves motion learning with semantic-structural feature-space supervision. We also establish PhyParam-Bench, a benchmark for physical-law consistency in image-to-video generation, with a multi-level protocol evaluating temporal dynamics, spatial stability, and semantic--physical alignment. Experiments show that PhyParam improves physical consistency while maintaining high visual fidelity, advancing explicit rigid-body physical-parameter control for image-to-video generation. We will publicly release the dataset, benchmark, and code to support future research.

Comments23 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑