arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07258cs.CVcs.AI

MV-STRIDE:通过分层能力建模使多模态大语言模型掌握多视角空间推理

MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

Jin Xu, Xiaojian Huang, Zhuodong Luo, Zhihong Zhang, Xin Liu, Jiansheng Wei, Xinzhi Wang, Jie Zhao, Xuejin Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态大模型多视角空间推理瓶颈,提出MV-STRIDE分层数据集与多阶段训练框架,显式建模能力依赖,在MMSI-Bench等基准上取得最优性能。

中文摘要 AI 辅助

尽管多模态大语言模型(MLLMs)在二维视觉语言任务上取得了快速进展,但由于现有数据集缺乏结构化的三维认知路径,稳健的多视角空间推理仍然是其根本性瓶颈。为解决这一问题,我们提出了MV-STRIDE,一个具有相互依赖和分解能力的多视角分层空间推理数据集。MV-STRIDE超越了扁平数据结构,显式建模了基础感知、场景理解和复杂上下文推理之间的依赖关系,提供了与人类空间认知一致的学习路径。我们开发了一个系统化的问答生成流程,利用多样化的三维场景源,强制执行跨视角依赖约束以防止单视角可解性,生成了多层级空间推理任务,并由基于认知的思维链监督支持复杂推理。大量评估表明,基于我们分层数据集的多阶段训练框架在多个空间推理基准上取得了最先进的性能,特别是在面向多视角的MMSI-Bench上。我们的方法使MLLMs能够在不同视角下保持稳健、三维一致的空间推理。代码和数据集可在以下网址获取:https URL。

英文摘要

Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust multi-view spatial reasoning remains a fundamental bottleneck due to the lack of structured 3D cognitive pathways in existing datasets. To address this, we introduce MV-STRIDE, a Multi-View hierarchical SpaTial Reasoning dataset with Interdependent and DEcomposed capabilitiEs. Moving beyond flat data structures, MV-STRIDE explicitly models the dependency relationships between foundational perception, scene understanding, and complex contextual reasoning, providing a coherent learning pathway aligned with human spatial cognition. We develop a systematic QA generation pipeline leveraging diverse 3D scene sources that enforces cross-view dependency constraints to prevent single-view solvability, generating multi-level spatial reasoning tasks supported by cognitively grounded chain-of-thought supervision for complex inference. Extensive evaluations demonstrate that our multi-stage training framework based on our hierarchical dataset achieves state-of-the-art performance across multiple spatial reasoning benchmarks, notably the multi-view oriented MMSI-Bench. Our approach enables MLLMs to maintain robust, 3D-consistent spatial reasoning across diverse viewpoints. The code and dataset are available at https://co1dspring.github.io/MV-STRIDE/.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Huawei Noah’s Ark Lab(华为诺亚方舟实验室)
  • Hefei University of Technology(合肥工业大学)

机构由 AI 辅助整理,请以论文原文为准。

↑