arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VerNav:优先验证器的低延迟视觉语言导航

VerNav: Verifier-First Low-Latency Vision-and-Language Navigation

Zhixin Wang, Chengzheyi Yao, Leyuan Liu, Xiaosong Zhang, Yongzhao Zhang

arXiv 2609.00920首次发表:更新:

发表机构

University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLN自回归生成延迟高的问题,提出VerNav优先验证器框架,通过批量动作验证与自适应生成器减少延迟,结合两阶段对齐方案提升性能,在R2R基准上每步LLM延迟降超10倍且性能具竞争力。

AI 中文摘要

视觉语言导航(Vision-and-Language Navigation, VLN)要求智能体根据自然语言指令在未见过的3D环境中导航。显式推理可提升指令理解与语义接地,但每一步的自回归生成会在多步导航中累积大量决策阶段延迟。我们提出VerNav,一种基于大语言模型(LLM)的低延迟VLN优先验证器框架。该验证器通过用批量动作验证替代每步自回归生成以减少决策阶段延迟,同时仅对不确定决策调用基于熵的自适应生成器以生成紧凑状态证据。为进一步提升验证器的导航性能,我们引入两阶段对齐方案:(i)VPO在静态验证器训练中改进局部动作偏好对齐;(ii)步骤级强化微调在动态任务执行期间,为多步导航回传提供密集进度奖励。在Room-to-Room(R2R)基准上的实验表明,VerNav仅验证器的决策路径在代表性基于LLM的VLN智能体中实现了有竞争力的导航性能,同时与自回归方法相比,每步平均决策阶段LLM延迟降低超过10倍。

英文摘要

Vision-and-Language Navigation (VLN) requires an agent to navigate through unseen 3D environments according to natural-language instructions. Explicit reasoning can improve instruction understanding and semantic grounding, but autoregressive generation at every step accumulates large decision-stage latency over multi-step navigation. We propose VerNav, a verifier-first framework for low-latency LLM-based VLN. The verifier reduces decision-stage latency by replacing per-step autoregressive generation with batched action verification, while an entropy-based adaptive generator is invoked only for uncertain decisions to produce compact state evidence. To further improve navigation performance with the verifier, we introduce a two-stage alignment scheme: (i) VPO improves local action-preference alignment in static verifier training, and (ii) step-level reinforcement fine-tuning provides dense progress rewards over multi-step navigation rollouts during dynamic task execution. Experiments on the Room-to-Room (R2R) benchmark show that the verifier-only decision path of VerNav achieves competitive navigation performance among representative LLM-based VLN agents while reducing average decision-stage LLM latency per step by more than $10\times$ compared with autoregressive methods.

Comments9 pages, 7 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑