arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WM-VLM:探求用于交错视觉-文本推理的内部世界模型

WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning

Yuheng Zha, Yilei Wang, Qiyue Gao, Junrong Chen, Yujia Wu, Zhengfeng Lai, Zhengzhong Liu, Eric P. Xing

arXiv 2609.34826首次发表:更新:

发表机构

UC San Diego; Institute of Foundation Models, MBZUAI; Carnegie Mellon University(加州大学圣迭戈分校; 基础模型研究所,穆罕默德·本·扎耶德人工智能大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出WM-VLM,通过轻量级世界模型生成中间视觉状态,使视觉-语言模型能结合文本与视觉进行空间推理,在2D和3D心理旋转任务上显著提升性能,最高达39.25个百分点。

AI 中文摘要

人类通常通过在心理上模拟视觉变换来解决空间问题。相比之下,传统的视觉-语言模型(VLM)主要通过语言进行推理。我们研究VLM是否能够通过同时使用文本和生成的视觉状态进行推理来解决空间问题。为此,我们提出了WM-VLM,它为预训练的VLM配备了一个轻量级的世界模型分支,用于生成中间视觉状态。我们的两阶段训练首先教会模型生成下一个视觉状态,然后利用该状态进行推理。我们通过编程方式构建具有可验证中间视觉状态的空间推理任务。这些任务使我们能够评估模型生成视觉状态的能力以及它在多大程度上依赖这些状态来回答问题。在2D和3D心理旋转任务上,WM-VLM始终优于监督微调的骨干模型,提升幅度最高达39.25个百分点。消融实验表明,这些提升依赖于生成的视觉状态,因为移除或破坏这些状态会显著降低性能。综合这些结果,内部世界模型为VLM在语言和视觉空间中同时推理提供了一条有前景的路径。

英文摘要

Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. Our two-stage training first teaches the model to generate the next visual state and then to use that state for reasoning. We programmatically construct spatial reasoning tasks with verifiable intermediate visual states. These tasks allow us to evaluate how well the model generates visual states and how much it relies on them to answer the question. On 2D and 3D mental rotation tasks, WM-VLM consistently outperforms the supervised fine-tuned backbone, with gains of up to 39.25 percentage points. Ablations suggest that these gains depend on the generated visual states, as removing or corrupting them sharply reduces performance. Together, these results suggest that internal world models offer a promising path toward VLMs that reason in both language and visual space.

Comments21 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑