arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PGN:基于盘古多模态基础模型的视觉语言导航系统的设计与实现

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model

Li Xian, Mingxi Li, Yizheng Wang, Yiming Shen, Qi Chen, Zhuoling Xiao

arXiv 2607.17806首次发表:更新:

发表机构

University of Electronic Science and Technology of China; School of Information and Communication Engineering(电子科技大学; 信息与通信工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于盘古多模态基础模型设计实现视觉语言导航系统PGN,训练分两阶段,先对齐视觉与语言编码器,再使模型适应专家轨迹,结合多种计算方式,在特定评估下取得一定指标,量化了离线专家动作对齐情况。

AI 中文摘要

视觉语言导航(VLN)要求具身智能体解释自然语言指令并根据时间顺序排列的视觉观察预测动作。将多模态大语言模型应用于VLN需要视觉语言对齐、紧凑的时间输入、动作空间基础以及在目标硬件上进行稳定训练。本技术报告介绍了PGN(盘古导航器),这是一个基于OpenPangu-7B构建的离线VLN动作预测系统。训练分两个阶段进行。首先,PGMM通过训练Q-Former和两层MLP投影仪,将冻结的EVA-ViT-G/14视觉编码器与冻结的语言主干对齐。其次,PGN使用五个观察窗口、依赖时期的时间采样和推理然后动作的输出格式,使对齐模型适应专家导航轨迹;此阶段冻结对齐的视觉通路并更新三个结构令牌嵌入和LoRA适配器。该实现结合了混合精度计算、选择性FP32计算和在八个Ascend 910B NPU上的DeepSpeed ZeRO-2。在对500条保留的专家轨迹进行教师强制、开环评估时,V9报告归一化动作匹配(NAM)为62.29%,非空率(NER)为100.00%。这些指标量化离线专家动作对齐而非闭环导航成功;评估误差积累、路径效率和目标完成情况仍是未来工作。

英文摘要

Vision-Language Navigation (VLN) requires an embodied agent to interpret a natural-language instruction and predict actions from temporally ordered visual observations. Adapting a multimodal large language model to VLN requires visual-language alignment, compact temporal inputs, action-space grounding, and stable training on the target hardware. This technical report presents PGN (Pangu Navigator), an offline VLN action-prediction system built on OpenPangu-7B. Training proceeds in two stages. First, PGMM aligns a frozen EVA-ViT-G/14 vision encoder with the frozen language backbone by training a Q-Former and a two-layer MLP projector. Second, PGN adapts the aligned model to expert navigation trajectories using five-observation windows, epoch-dependent temporal sampling, and a reasoning-then-action output format; this stage freezes the aligned visual pathway and updates three structural-token embeddings and LoRA adapters. The implementation combines mixed-precision computation, selective FP32 computation, and DeepSpeed ZeRO-2 on eight Ascend 910B NPUs. Under teacher-forced, open-loop evaluation on 500 held-out expert trajectories, V9 reports a 62.29% Normalized Action Match (NAM) and a 100.00% Non-empty Rate (NER). These metrics quantify offline expert-action alignment rather than closed-loop navigation success; evaluating error accumulation, path efficiency, and goal completion remains future work.

Comments6 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑