arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34687cs.AI

VCN-Bench:面向先验视觉经验空间推理的视频情境化导航基准

VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience

Siqi Zhang, Meng Wei, Chenyang Wan, Shaohao Zhu, Shufan Shen, Xihui Liu, Zhihua Wei, Tai Wang, Jiangmiao Pang

首次发表
浏览论文内容

中文总结 AI 辅助

VCN-Bench是一个视频情境化导航基准,用于评估多模态大语言模型基于先验视频进行闭环空间推理的能力,实验发现导航性能有限且存在目的地解析与导航间的显著差距。

中文摘要 AI 辅助

空间推理是具身智能体的基础能力,然而,空间理解能否被延续以指导顺序交互仍不明确。现有的空间推理基准通常止步于离线预测,而导航基准则将空间推理作为指令跟随和探索的一部分进行评估。我们提出了VCN-Bench,一个视频情境化导航基准,用于探测多模态大语言模型(MLLMs)中基于先验视觉经验的闭环空间推理。给定一段涵盖初始位置和目的地的先验视频,智能体需要推理出指令指定的目标,并利用推断出的空间情境导航至该目标。VCN-Bench基于Matterport3D构建,包含五种指令类型、10万个训练回合和1250个评估回合。导航作为主要评估方式,而诊断性目标识别有助于区分目的地解析错误与后续导航失败。我们进一步提出了MV-DualVLN,一个规划导向的基线方法,该方法联合利用先验视频和回合内观测。实验揭示了有限的导航性能、显著的目的地解析与导航之间的差距,以及即使在正确识别目的地后仍频繁出现的导航失败。

英文摘要

Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning as part of instruction following and exploration. We introduce VCN-Bench, a \textbf{V}ideo-\textbf{C}ontextualized \textbf{N}avigation benchmark for probing closed-loop spatial reasoning over prior visual experience in MLLMs. Given a prior video covering both the initial location and destination, the agent is tasked with reasoning out the instruction-specified target and navigating toward it with the inferred spatial context. Built on Matterport3D, VCN-Bench contains five instruction types, 100k training episodes, and 1,250 evaluation episodes. Navigation serves as the primary evaluation, while diagnostic goal identification helps distinguish destination-resolution errors from subsequent navigation failures. We further propose MV-DualVLN, a planning-oriented baseline that jointly leverages prior video and in-episode observations. Experiments reveal limited navigation performance, a substantial destination-resolution-to-navigation gap, and frequent navigation failures even after correct destination identification.

发表机构

  • Tongji University(同济大学)
  • Shanghai AI Laboratory(上海人工智能实验室)
  • The University of Hong Kong(香港大学)
  • Zhejiang University(浙江大学)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑