arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

世界嵌入基准

World Embedding Benchmark

Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li, Fengyu Cai, Hou Pong Chan, Jialin Yu, Hao Zhang, Chenghua Lin, Chenghao Xiao

arXiv 2610.03632首次发表:更新:

AI 中文总结

本文提出世界嵌入基准,含8,000个物理模拟视频案例,评估跨模态对齐与物理信息可恢复性,并揭示对齐与回归的权衡,且检索增强可提升视频生成的物理保真度。

AI 中文摘要

物理保真度在世界模型和视频生成中受到越来越多的关注,然而视频表示如何编码物理信息仍未被充分理解。我们引入了世界嵌入基准,包含来自80个族系的8,000个受控模拟案例,涵盖流体力学、固体力学、动力学以及光学与电磁学。每个案例将渲染视频与模拟衍生的物理注释配对,支持三个互补任务:文本-视频检索、物理属性回归和视频-描述对的多选分类。我们利用这些任务来区分跨模态物理对齐与定量物理信息的可恢复性。评估的预训练全模态嵌入模型显示出较弱的检索性能和接近偶然水平的族内对分类,而轻量级探针能从冻结的视频嵌入中恢复有用的物理信息。使用物理特定的视频-文本对进行持续对比训练改善了检索和族内对分类,但降低了物理属性回归性能,揭示了对齐与定量信息可恢复性之间的权衡。最后,我们利用嵌入检索参考视频,用于与MiniMax-H3的检索增强生成。检索到的参考提高了生成视频的物理保真度,在我们的实验中,更强的检索模型带来了更大的增益。综上所述,这些发现强调了联合评估物理对齐和属性可恢复性的必要性,并展示了物理表示在改善视频生成方面的实用性。

英文摘要

Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑