发表机构
The University of Tokyo; Carnegie Mellon University(东京大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出情境观察者锚定概念,构建POVBench基准,评估视觉-语言模型在具身任务中从情境线索推断说话者视角并定位目标的能力,发现明确分解空间推理可提升定位性能。
AI 中文摘要
在机器人等具身任务中,对语言指令进行推理通常需要从说话者的情境化视角理解空间关系。人类从共享的环境知识、活动情境和常识中推断出这种视角。最近的视觉-语言模型(VLMs)似乎具备空间推理能力,但它们从情境线索中推断说话者视角,并从该视角解释情境化空间关系的能力仍不清楚。我们将这种能力称为情境观察者锚定。为了研究这种能力,我们构建了视角基准(POVBench),这是一个包含3D场景和查询的数据集,在自然具身交流中区分了推断、陈述和给定三种形式的观察者锚定。给定多视角观察和自然语言句子,模型必须从情境化的空间和情境线索中定位未见过的或欠指定的目标。在最先进的VLMs中,即使观察者锚定被明确给出,从方向性语言中定位目标仍然具有挑战性。我们发现,对观察者相对空间推理的明确分解能提高目标定位性能。我们的项目页面可在该https URL获取。
英文摘要
Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated perspective. Humans infer such perspectives from shared environmental knowledge, activity context, and commonsense. Recent vision-language models (VLMs) appear capable of spatial reasoning, but their ability to infer a speaker's viewpoint from contextual cues and interpret situated spatial relations from that viewpoint remains unclear. We call this capability contextual observer grounding. To study this capability, we construct the Point-of-View Benchmark (POVBench), a dataset of 3D scenes and queries that disentangles Inferred, Stated, and Given forms of observer grounding in natural embodied communication. Given multi-view observations and a natural-language sentence, models must localize unseen or underspecified targets from situated spatial and contextual cues. Across state-of-the-art VLMs, localizing targets from directional language remains challenging, even when observer grounding is made explicit. We find that explicit breakdowns of observer-relative spatial reasoning improve target localization. Our project page is available at https://mimo-owl.github.io/POVBench/.
CommentsAccepted to EMNLP 2026 Findings