发表机构
Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究审计多模态世界模型,发现其视频与文本预测间存在内部错位及与真实物理环境的外部错位,通过契约阶梯方法在四种机制下测试,揭示统一骨干难以同时实现正确推理、内部一致与物理保真。
AI 中文摘要
世界模型是一种根据当前环境条件生成接下来会发生什么的系统,正越来越多地以多模态生成为目标进行实现。然而,同时生成多种模态,例如视觉模拟与以文本形式呈现的物理状态预测,会引入跨模态不一致的风险。分别测试时,两种输出可能看起来都令人信服,但仍然存在分歧:模型可以在一种模态中计算出球应该反弹,然后在另一种模态中生成不反弹的结果,更不用说完全偏离真实世界动力学了。在这项工作中,我们明确关注两种失败:\n内部错位,即世界模型生成的视频与同一世界模型在不同模态中的预测之间的分歧;以及外部错位,即世界模型的生成与解析物理环境之间的分歧。我们推导出事件、幅度、时序的通用契约,并构建一个基于物理的流水线,使比较在外部和内部设置中均可测量。然后,我们询问逐步提供模型自身的契约(内部设置的A阶梯)或修正后的物理契约(外部设置的B阶梯)是否能弥合相应的差距。在四种机制和20种设置中,我们发现,虽然语言模态相对于真实环境正确回答了所有22个文本探针,但中性视频常常存在分歧,这表明当前统一的骨干网络可能无法同时具备正确的推理、内部一致性和外部物理保真度。
英文摘要
World models, systems that generate what happens next given current environmental conditions, are increasingly being implemented with multi-modal generation in mind. However, generating multiple modalities simultaneously, such as visual simulations alongside physical state predictions in the form of text, introduces the risk of cross-modal inconsistency. Tested separately, both outputs may look convincing while still disagreeing: a model can calculate that a ball should rebound in one modality, then generate no rebound in another modality, to say nothing of diverging from real-world dynamics entirely. In this work we focus on two failures explicitly: \emph{Internal misalignment}, the disagreement between the world model's generated video and the same world model's prediction in a different modalities, and \emph{external misalignment} the disagreement between the world model's generation and an analytic physical environment. We derive common contracts of event, magnitude, timing, and construct a physics grounded pipeline to make comparisons measurable in both external and internal settings. We then ask whether progressively supplying the model's own contract (the A ladder for the internal setting) or a corrected physical contract (the B ladder for the external setting) closes the respective gaps. Across four mechanisms and 20 settings, we find that while language answers all 22 text probes correctly with respect to the true environment, the neutral video is often in disagreement, suggesting that the current unified backbones may not be capable of correct reasoning, internal consistency, and external physical fidelity all at once.