arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HiPhy:面向物理合理多原理视频生成的层次对齐方法

HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag

arXiv 2610.02197首次发表:更新:

发表机构

Virginia Tech; Qualcomm AI Research(弗吉尼亚理工大学; 高通人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HiPhy提出层次物理对齐强化学习框架,通过局部与全局双层目标解决视频生成中多物理原理协同问题,并构建50K数据集和MultiPhyBench基准,显著提升物理常识与语义对齐。

AI 中文摘要

视频生成模型已实现了显著的视觉保真度,并具有成为通用世界模拟器的巨大潜力。尽管取得了这些进展,它们仍然无法生成符合物理定律的视频。在现实场景中,当多个物理原理必须在同一视频中协同作用时,这一问题变得更加明显;例如,“气球向上漂浮,同时蒸汽从锅中升起”需要浮力和流体动力学同时且连贯地展开。然而,现有方法大多忽略多原理交互,仅关注每个视频中的单一原理。我们提出了HiPhy(层次物理对齐),一种强化学习框架,通过双层目标将视频生成锚定在物理定律上:局部强制单个物理原理的时间动态,全局确保整个场景的物理和语义连贯性。为支持多原理生成,我们构建了一个包含50K提示的数据集,并引入了一个提示基准MultiPhyBench,涵盖多种共现物理事件。实验表明,HiPhy显著优于先前方法和基线,在各种基准上大幅提升了物理常识和语义对齐,其中在涉及多个并发物理原理的场景中提升最大,而竞争方法在这些场景中性能下降最为严重。

英文摘要

Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.

CommentsProject page: https://hiphy-video.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑