ConsiSpace:学习几何一致性对视频空间推理很重要
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
浏览论文内容
中文总结 AI 辅助
针对视频空间推理中现有模型难以聚合一致空间证据的问题,提出ConsiSpace框架,通过构建几何一致内存及采用统一一致性自监督强化学习,在三个空间推理基准测试上取得成绩提升,平均得分比最强基线高12.6分。
中文摘要 AI 辅助
视频空间推理对导航感知和长视频问答至关重要,模型需在变化视角下推断长距离空间关系。但现有多模态大语言模型以语义为中心,难以可靠聚合冗余视频观测中的一致空间证据。为此提出ConsiSpace,一个几何一致性感知框架,将空间一致性转化为证据组织原则和明确的后SFT学习信号。构建几何一致内存,利用高效组织策略保存空间证据,通过统一一致性自监督强化学习提升跨视图稳定性。在三个基准测试上实验显示成绩提升,比最强基线平均得分提高12.6分。
英文摘要
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.
发表机构
- School of Intelligent Science and Technology, Nanjing University(南京大学智能科学与技术学院)
- School of Computer Science, Peking University(北京大学计算机科学学院)
- Beijing Academy of Artificial Intelligence(北京人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。