发表机构
CSIRO Technology(联邦科学与工业研究组织)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探索基于大型预训练模型的视觉机器人导航策略新方法,提出GemNav,仅对语言塔用LoRA适配冻结的多模态大语言模型用于中短程航点导航,无辅助视觉编码器等,在小语料库上零样本转移到未见环境,提供了数据高效、可部署的导航替代方案。
AI 中文摘要
基于大型预训练模型构建的视觉导航策略通常采用专用视觉编码器、定制动作头并在数千小时的跨实体数据集上训练。本文提出疑问,这一方法是否必要。为此引入GemNav,它仅对语言塔使用低秩适应(LoRA)来适配冻结的多模态大语言模型,用于中短程航点导航,无辅助视觉编码器和连续回归头。航点和分类导航信号共享由语言模型头生成的单个离散令牌词汇表,软解码辅助损失恢复纯交叉熵训练丢弃的度量结构。在单个8.7小时开放语料库上,该策略零样本转移到四个未见环境,在20次真实世界试验中距目标0.25 - 0.42米内停止。基于短图像历史进行条件设定可改善离线指标,但对机器人无益处。结果表明冻结的多模态大语言模型的离散令牌适应可为基础模型机器人导航提供数据高效、可部署的替代方案。
英文摘要
Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether this recipe is necessary. In this paper, we introduce GemNav, a visual robot navigation policy that adapts a frozen Multimodal Large Language Model (MLLM) for short-to-medium horizon waypoint navigation using Low-Rank Adaptation (LoRA) on the language tower alone, with no auxiliary visual encoder and no continuous regression head. Waypoints and categorical navigation signals share a single discrete token vocabulary generated by the language-model head, and a soft-decoded auxiliary loss recovers the metric structure that pure cross-entropy training discards. On a single 8.7-hour open corpus, roughly three orders of magnitude smaller than competing training sets, the policy transfers zero-shot to four physically distinct unseen environments and stops within 0.25-0.42m of the goal across 20 real-world trials covering an open carpark, an obstacle carpark, a long outdoor chemical yard, and an indoor warehouse. Conditioning on short image histories improves offline metrics but yields no robot benefit, pointing to a ceiling on what temporal context adds once pretrained vision features are in place. These results indicate that discrete-token adaptation of frozen MLLMs can provide a data-efficient, deployable alternative for foundation model robot navigation.