arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39195cs.CV

Uruqi:从视觉经验中学习空间认知

Uruqi: Learning Spatial Cognition from Visual Experience

Shichao Li, Meiqi Wang, Fei Su, Zhicheng Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出URUQI方法,通过合成连续视觉经验训练VLM,提升自运动跟踪和世界映射能力,在Uruqi基准上准确率从15.84%提升至47.73%,并优于基线模型。

中文摘要 AI 辅助

空间智能要求在具身智能体移动时保持对世界的一致理解。与人类一样,智能体必须利用自身的运动来解释观测之间的变化,并相应地更新物体位置和空间关系。尽管空间后训练已大幅拓展了视觉语言模型(VLM)的空间智能,但它们在两项原子空间能力上仍存在困难:跟踪自身运动以及在运动过程中绘制周围世界的地图。为解决这一差距,我们在每个训练回合内对交错的原子能力提供密集的多轮监督,模拟一个连续移动的智能体在观察时进行推理的视觉经验。为扩大规模,我们在广泛的3D场景中合成了11,738条由动机驱动的相机轨迹,支持在每次视觉经验中进行自运动跟踪、持久物体映射和丰富的空间操作。通过训练模型对这些原子问题进行推理,我们的URUQI$_{\mathrm{Syn}}$-8B在我们的Uruqi基准(包含2.7k个回合中的52k个问题)上将准确率从15.84%提升至47.73%。URUQI-SI-Mix-8B进一步达到50.41%,与GPT-6 Astra取得的50.08%相当。仅使用我们合成的数据进行训练,URUQI$_{\mathrm{Syn}}$-8B在三个外部空间基准上相对于其InternVL3-8B骨干网络实现了平均相对准确率提升17.13%。这些结果凸显了连续视觉经验作为开发VLM空间认知的可扩展监督来源。

英文摘要

Spatial intelligence requires maintaining a coherent understanding of the world as the embodied agent moves. Like humans, the agent must use its own motion to interpret changes across observations and update object locations and spatial relations accordingly. Despite spatial post-training having substantially broadened the spatial intelligence of vision-language models (VLMs), they still struggle with two atomic spatial capabilities: tracking self-motion and mapping the surrounding world during motion. To address this gap, we provide dense multi-turn supervision over interleaved atomic capabilities within each training episode, mimicking the visual experience of a continuously moving agent that reasons as it observes. To scale this up, we synthesize 11,738 motif-driven camera trajectories over a broad range of 3D scenes, supporting self-motion tracking, persistent object mapping, and rich spatial operations within each visual experience. By training models to reason over these atomic questions, our URUQI$_{\mathrm{Syn}}$-8B improves accuracy from 15.84% to 47.73% on our Uruqi benchmark comprising 52k questions across 2.7k episodes. URUQI-SI-Mix-8B further reaches 50.41%, comparable to the 50.08% achieved by GPT-6 Astra. Trained solely on our synthesized data, URUQI$_{\mathrm{Syn}}$-8B achieves an average relative accuracy improvement of 17.13% over its InternVL3-8B backbone across three external spatial benchmarks. These results highlight continuous visual experience as a scalable source of supervision for developing spatial cognition in VLMs.

发表机构

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑