arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36219cs.CV

LeRF:学习参考坐标系以进行视角推理

LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning

Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen, Yibo Yang, Bolin Lai, Junho Kim, James Matthew Rehg

首次发表
浏览论文内容

中文总结 AI 辅助

LeRF框架通过训练视觉语言模型构建显式参考坐标系,解决视角推理中默认相机视角的问题,并在多个基准上提升性能。

中文摘要 AI 辅助

视角推理是空间智能的基本组成部分,要求模型从指定视角(如另一实体或想象观察者的视角)解释空间关系。尽管视觉语言模型(VLMs)的空间推理能力不断增强,但它们在视角推理方面仍存在困难,当查询需要从不同视角进行推理时,往往默认采用相机视角。我们提出了用于视角推理的学习参考坐标系(LeRF)框架,该框架训练VLMs构建和使用显式参考坐标系进行依赖于视角的推理。给定图像和查询,LeRF决定是否需要坐标系。如果需要,它定位参考实体并预测坐标系的原点和以实体为中心的参考坐标系。一个轻量级渲染器将坐标系叠加到图像上,使后续推理能够基于这些视觉线索进行,而无需外部感知模型或显式3D重建。为了学习这一过程,我们首先进行监督微调以教授选择性工具调用和参考坐标系预测,随后在空间VQA对上进行强化学习以改进坐标系引导的推理。在多个视角推理基准上,LeRF持续优于其骨干模型,并与现有开源方法相比取得了强劲性能。进一步评估还显示了参考坐标系定位和方向估计的改进,支持了学习参考坐标系对视角依赖推理的有效性。

英文摘要

Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.

发表机构

  • Zhiyuan College, Shanghai Jiao Tong University(上海交通大学致远学院)
  • University of California, Merced(加州大学默塞德分校)
  • Shanghai Jiao Tong University(上海交通大学)
  • Amazon AGI(亚马逊AGI)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑