发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有空间智能基准无法评估综合导航能力的问题,提出EgoPathBench数据集与五项任务基准,测试九个视觉语言模型零样本航点决策,最高得分仅28.3,微调Qwen 3.5 4B后得分提升至38.9。
AI 中文摘要
零样本航点导航要求视觉语言模型从当前第一人称观察中选择一系列对智能体可行且能到达目标的空间动作序列,这对当今基础视觉语言模型的综合空间智能提出了联合要求。现有的空间智能基准主要评估关系、方向或目标的孤立判断,因此不能直接衡量结合目标识别、动作后果评估、距离估计和路径规划所需的综合导航能力。为填补这一评估空白,我们引入了EgoPathBench,一个用于第一人称航点决策的数据集和五项任务基准。每个问题呈现一张自我中心RGB图像、一个自然语言目标和编号的可见航点;模型返回可通行的候选点或有序路线。预测结果根据候选可行性、相邻边合法性和在点智能体或具身几何下的目标到达情况进行评估。EgoPathBench包含31,852个训练问题、1,345个验证问题和1,111个基准问题,并且每个路线问题至少保留一条经过几何验证的参考路线。在九个视觉语言模型中,最高的EgoPath得分仅为28.3。排名最高的模型在点路径上达到35.9%的成功率,但在具身路径和意图路径上分别仅为2.9%和4.0%,表明当前模型在具身约束下形成完整、目标一致的路线方面仍然有限。除评估数据外,我们还发布了相应的训练资源。在发布的训练集上对Qwen 3.5 4B进行微调,将其EgoPath得分从3.9提升至38.9,并在三个外部空间基准上改善了所有四项报告的评估,增益为1.4至9.6个百分点。
英文摘要
Zero-shot waypoint navigation requires vision-language models to select, from the current first-person observation, a sequence of spatial actions that is feasible for the agent and reaches the goal, placing joint demands on the integrated spatial intelligence of today's foundation VLMs. Existing spatial-intelligence benchmarks primarily evaluate isolated judgments of relations, directions, or targets and therefore do not directly measure the integrated navigation ability required to combine target recognition, action-consequence assessment, distance estimation, and path planning. To fill this evaluation gap, we introduce EgoPathBench, a dataset and five-task benchmark for first-person waypoint decision-making. Each question presents an egocentric RGB image, a natural-language goal, and numbered visible waypoints; a model returns traversable candidates or an ordered route. Predictions are evaluated for candidate feasibility, adjacent-edge legality, and goal arrival under point-agent or embodied geometry. EgoPathBench contains 31,852 training, 1,345 validation, and 1,111 benchmark questions and retains at least one geometrically verified reference route for every route question. Across nine VLMs, the highest EgoPath Score is only 28.3. The top-ranked model reaches 35.9% success on Point Path, but only 2.9% and 4.0% on Embodied Path and Intent Path, respectively, showing that current models remain limited in forming complete, goal-consistent routes under embodiment constraints. Beyond the evaluation data, we release the corresponding training resource. Fine-tuning Qwen 3.5 4B on the released training split raises its EgoPath Score from 3.9 to 38.9 and improves all four reported evaluations across three external spatial benchmarks, with gains of 1.4--9.6 points.
Comments18 pages, including supplementary material