arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Physis-Lang:作为视频世界模型的物理表示的自演化语言

Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model

Liming Lu, Xianzheng Ma, Wenkun He, Guanqi Zhan, Yilin Zhao, Junyu Chen, Mengyao Xu, Jiaojiao Fan, Wenhang Ge, Yuchao Gu, Yunze Liu, Boyi Li, Zhen Dong, Victor Prisacariu, Ming-Yu Liu, Song Han, Han Cai

arXiv 2609.40358首次发表:更新:

发表机构

NVIDIA; University of Oxford; MIT(英伟达; 牛津大学; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Physis-Lang通过自演化语言表示物理过程,结合PhysCapBench和智能体循环优化,提升视频世界模型的物理合理性,并超越专有模型。

AI 中文摘要

视频世界模型预期能够预测物理世界如何演化,然而它们常常生成在视觉上看似合理但违反基本物理原理的视频。现有方法通常假设自然语言不足以表示可靠生成所需的物理知识,因此引入额外的视觉、潜在、数值或基于规划的信号。我们重新审视这一假设,并引入Physis-Lang,一个自演化框架,将物理语言视为跨数据整理、模型训练和视频生成的共享且可优化的表示。Physis-Lang通过描述相关实体、原因、交互、支配原理、时间演化及效应的语言来表示物理过程。为了改进这一表示,我们构建了PhysCapBench,将物理过程分解为原子断言,并使用召回率和精确率评估描述。一个智能体循环迭代分析断言级错误,并优化用于生成物理描述指令。Physis-Lang进一步将模型缺陷转化为文本描述,并使用语言引导的检索来识别覆盖缺失物理过程的视觉多样视频。在四个广泛使用的物理视频基准上,使用Wan和Cosmos骨干网络的实验表明,物理合理性得到了一致的提升。值得注意的是,从开源的Cosmos3-Nano骨干网络出发,我们Physis-Lang增强的模型超越了领先的专有Veo 3.1模型。

英文摘要

Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefore introduce additional visual, latent, numerical, or planning-based signals. We revisit this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. Physis-Lang represents physical processes through language that describes their relevant entities, causes, interactions, governing principles, temporal evolution, and effects. To improve this representation, we construct PhysCapBench, which decomposes physical processes into atomic assertions and evaluates captions using recall and precision. An agentic loop iteratively analyzes assertion-level errors and refines the instruction used to produce physical captions. Physis-Lang further converts model deficiencies into textual descriptions and uses language-guided retrieval to identify visually diverse videos that cover missing physical processes. Experiments on four widely used physical video benchmarks with Wan and Cosmos backbones demonstrate consistent improvements in physical plausibility. Notably, starting from open-source Cosmos3-Nano backbones, our Physis-Lang-enhanced models surpass the leading proprietary Veo 3.1 model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑