arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17657cs.AIcs.CVcs.MM

OrientSAM:通过方向感知空间对齐减轻多模态空间推理中以相机为中心的捷径

OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment

Wenxiao Fan, Hang Yin, Kan Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态模型空间推理中以相机为中心的问题,提出OrientSAM框架,通过方向感知令牌、傅里叶角度编码及课程学习策略注入方向信息并改善推理,构建数据管道生成监督,实验证明其在多任务中表现出色,能减轻捷径行为,实现更强大的以对象为中心的空间推理。

中文摘要 AI 辅助

多模态大语言模型在需要视角转换的空间推理方面仍存在困难。它们常依赖以相机为中心的线索而非从参考对象视角进行推理,在非相机参考设置中导致系统性错误。本文首先分析了这种失败模式,表明对象方向是这种以相机为中心的捷径行为的关键因素。为解决此问题,提出OrientSAM,一个用于多模态模型的方向感知空间对齐框架。它通过方向感知令牌和基于傅里叶的角度编码将明确的方向信息注入多模态表示,并采用课程学习策略逐步改善视角感知推理。还构建了空间数据构建管道以从大规模图像生成方向感知空间监督。在多个数据集上的实验表明OrientSAM始终优于强大的基线,尤其在非相机视角、以人为中心和方向敏感任务上。结果进一步证明明确的方向建模对于减轻以相机为中心的捷径行为和在多模态模型中实现更强大的以对象为中心的空间推理很重要。

英文摘要

Multimodal large language models (MLLMs) still struggle with spatial reasoning that requires perspective transformation. In particular, they often rely on camera-centric cues rather than reasoning from the reference object's viewpoint, leading to systematic errors in non-camera reference settings. In this paper, we first analyze this failure mode and show that object orientation is a key factor underlying such camera-centric shortcut behavior. To address this issue, we propose OrientSAM, an orientation-aware spatial alignment framework for multimodal models. OrientSAM injects explicit orientation information into multimodal representations through orientation-aware tokens and Fourier-based angle encoding, and further adopts a curriculum learning strategy to progressively improve perspective-aware reasoning. In addition, we build a spatial data construction pipeline to generate orientation-aware spatial supervision from large-scale images. Experiments on Spatial-MM, ViewSpatial, and 3DSRBench show that OrientSAM consistently outperforms strong baselines, especially on non-camera-view, person-centric, and orientation-sensitive tasks. The results further demonstrate that explicit orientation modeling is important for mitigating camera-centric shortcut behavior and enabling more robust allocentric spatial reasoning in multimodal models.

↑