arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WOVEN:将视觉世界建模融入多模态大语言模型

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, Chenlei Guo, Eric P. Xing, Mohit Bansal, Manling Li

arXiv 2610.12417首次发表:更新:

发表机构

Northwestern University; Carnegie Mellon University; UNC Chapel Hill; Amazon(西北大学; 卡内基梅隆大学; 北卡罗来纳大学教堂山分校; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多模态大语言模型的视觉推理缺陷,推出WOVEN训练源与基准,发现其可提升模型在多数外部基准的性能,确立视觉过渡推理为视觉世界建模的可复用基础。

AI 中文摘要

多模态大语言模型(MLLMs)在空间、具身、物理及时间推理方面存在不足。我们假设这些缺陷反映了视觉过渡推理的共同缺失,并测试该能力是否可作为通用训练基元,让不同模型从不同监督源学习并跨任务复用,同时辅以系统的训练方案。现有基准分别记录了这些缺陷,但不支持跨场景、动作和推理操作的受控比较。因此,我们推出WOVEN,这是一个用于视觉过渡推理的训练源和基准,它按场景、动作和推理类型组织过渡监督,使用视频预训练生成模型的多样化、真实回滚数据:共36076个示例,涵盖20种场景类型、5种动作类型和8种推理类型。我们首先评估了38个前沿MLLMs(如GPT-5.4和Qwen3-VL-235B-A22B),发现存在显著且系统性的缺陷:即使是最强的模型也远低于人类水平,且这些缺陷在不同模型家族中反复出现,随模型规模增大仍持续存在。随后,我们在WOVEN上训练了不同规模的MLLMs,发现它们学习到了可广泛迁移的通用能力:仅约2000个示例的训练子集,就能共同提升26个外部基准中的22个,提升幅度最高达27.3个百分点;WOVEN数据可替代任务自身训练数据的30%-50%,且准确率相当。受控比较进一步得出了视觉世界建模的训练方案,已在保留基准上得到前瞻性验证:选择监督时应依据其教授的推理操作,而非所展示的动作、场景或领域;为提升鲁棒性,应优先选择视觉状态的更大变化。我们的研究确立了视觉过渡推理作为MLLMs系统性视觉世界建模训练的可复用基础。

英文摘要

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑