arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.01805cs.CV

SpaceR:强化多模态大语言模型的视频空间推理能力

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

  • National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机学院,北京大学)
  • Nanyang Technological University(南洋理工大学)
  • WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

机构由 AI 辅助整理,请以论文原文为准。

Kun Ouyang, Yuanxin Liu, Haoning Wu, Yi Liu, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun

更新

AI总结:

针对现有多模态大语言模型视频空间推理能力不足的问题,提出SpaceR框架,通过构建含15.1万样本的专用数据集和带地图想象机制的空间引导RLVR方法,在空间推理基准上超越GPT-4o,达到领先专有模型水平。

AI中文摘要:

视频空间推理指从观测到的视频帧中推断潜在的空间结构,这对现有的多模态大语言模型(MLLMs)构成了重大挑战。这一局限性主要源于两个方面:1)缺乏适用于该任务的高质量数据集;2)缺少能培养空间推理能力的有效训练策略。受可验证奖励强化学习(Reinforcement Learning with Verifiable Reward, RLVR)在解锁大语言模型推理能力方面取得成功的启发,本研究旨在通过RLVR范式提升多模态大语言模型的视频空间推理能力。为此,我们提出了SpaceR框架。首先,我们构建了SpaceR-151k数据集,其中包含9.1万个覆盖多种空间推理场景且答案可验证的问题,以及6万个用于维持通用多模态理解能力的样本。其次,我们提出了空间引导可验证奖励强化学习(Spatially-Guided RLVR, SG-RLVR),这是一种新型强化学习方法,它在组相对策略优化(Group Relative Policy Optimization, GRPO)的基础上引入了全新的地图想象机制,鼓励模型在思考过程中推断空间布局,从而实现更高效的空间推理。大量实验表明,SpaceR在空间推理基准(如VSI-Bench、STI-Bench和SPAR-Bench)上取得了当前最优性能,同时在视频理解基准(如Video-MME、TempCompass和LongVideoBench)上保持了具有竞争力的结果。值得注意的是,SpaceR在VSI-Bench上的准确率比先进的GPT-4o高出11.6%,与领先的专有模型Gemini-2.0-Flash性能相当,凸显了我们的SpaceR-151k数据集和SG-RLVR在强化多模态大语言模型空间推理能力方面的有效性。代码、模型和数据集可在https://github.com/OuyangKun10/SpaceR获取。

英文摘要:

Video spatial reasoning, which involves inferring the underlying spatial structure from observed video frames, poses a significant challenge for existing Multimodal Large Language Models (MLLMs). This limitation stems primarily from 1) the absence of high-quality datasets for this task, and 2) the lack of effective training strategies to develop spatial reasoning capabilities. Motivated by the success of Reinforcement Learning with Verifiable Reward (RLVR) in unlocking LLM reasoning abilities, this work aims to improve MLLMs in video spatial reasoning through the RLVR paradigm. To this end, we introduce the $\textbf{SpaceR}$ framework. First, we present $\textbf{SpaceR-151k}$, a dataset with 91k questions spanning diverse spatial reasoning scenarios with verifiable answers, and 60k samples for maintaining general multimodal understanding. Second, we propose $\textbf{Spatially-Guided RLVR (SG-RLVR)}$, a novel reinforcement learning approach that extends Group Relative Policy Optimization (GRPO) with a novel map imagination mechanism, which encourages the model to infer spatial layouts in the thinking process, thereby facilitating more effective spatial reasoning. Extensive experiments demonstrate that SpaceR achieves state-of-the-art performance on spatial reasoning benchmarks (e.g., VSI-Bench, STI-Bench, and SPAR-Bench), while maintaining competitive results on video understanding benchmarks (e.g., Video-MME, TempCompass, and LongVideoBench). Remarkably, SpaceR surpasses the advanced GPT-4o by 11.6\% accuracy on VSI-Bench and is on par with the leading proprietary model Gemini-2.0-Flash, highlighting the effectiveness of our SpaceR-151k dataset and SG-RLVR in reinforcing spatial reasoning ability of MLLMs. Code, model, and dataset are available at https://github.com/OuyangKun10/SpaceR.

↑