发表机构
Southern University of Science and Technology; Peng Cheng National Laboratory; Guangdong University of Technology(南方科技大学; 鹏城国家实验室; 广东工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究连续环境中视觉与语言导航问题,提出通过自回归轨迹生成在2D像素空间微调视觉语言模型来学习导航交互的方法,实验验证该方法能显著提升性能,旗舰模型在有限资源和数据下达最优水平。
AI 中文摘要
受益于大规模预训练数据中强大的先验知识和新兴的常识推理能力,大语言模型(LLMs)在许多研究领域展现出前所未有的泛化能力。近期,通过视觉语言模型(VLMs)将视觉嵌入投影到语言空间以实现模拟到真实和跨场景泛化,已成为连续环境中的视觉与语言导航(VLN-CE)领域的主流范式。VLN要求具身智能体根据自然语言指令在未见环境中导航。我们强调VLN任务可分解为一系列子任务,每个子任务对应与指令描述的环境进行3D空间交互的过程。然而,这种涉及沿深度感知方向进入图像的空间交互对主要在RGB图像对话上训练的VLMs来说颇具挑战。我们提出一种替代方法:通过自回归轨迹生成在2D像素空间中微调VLMs以直接学习导航交互。给定语言指令和历史观察,模型顺序预测一系列像素坐标,从当前观察的底部中心绘制轨迹。实验进一步验证像素空间轨迹监督显著提升VLN性能,且旗舰模型在相对有限计算资源和训练数据下达到了当前最优性能水平。
英文摘要
Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models (LLMs) have shown unprecedented generalization capabilities in many research fields. Recently, projecting visual embeddings into the language space via vision-language models (VLMs) to achieve sim-toreal and cross-scene generalization has become a prevailing paradigm in the field of Vision-and-Language Navigation in Continuous Environments (VLN-CE). VLN requires an embodied agent to navigate through unseen environments following natural linguistic instructions. We emphasize that a VLN task can be decomposed into a sequence of sub-tasks, each corresponding to a process of 3D spatial interaction with the environments described by instructions such as "walk to the end of the sofa and turn left." However, such spatial interactions involving moving into the image along the direction of depth sensing are puzzling for VLMs as they were predominantly trained on conversations with RGB images. Rather than incorporating depth or 3D geometric information-which VLMs rarely encounter during pretrainingwe propose an alternative approach: fine-tuning VLMs to learn navigation interactions directly in 2D pixel space through autoregressive trajectory generation. Given a linguistic instruction and historical observations, our model sequentially predicts a series of pixel coordinates, drawing a trajectory from the bottom center of the current observation. While prior work has proved that pixel-goal supervision outperforms learning of discrete actions, our experiments further verify that the supervision of pixel-space trajectory significantly enhances VLN performance. Moreover, we demonstrate that our flagship model achieves state-of-the-art level performance with relatively limited computational resources and training data.