arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于阿克曼转向移动机器人的连续环境视觉-语言导航的仿真到现实迁移

Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot

Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara

arXiv 2610.07192首次发表:更新:

发表机构

University of Moratuwa(莫拉图瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对传统VLN模型依赖导航图等限制,提出一种无需这些假设的连续环境VLN方法,通过跨模态注意力架构在仿真训练后利用真实数据微调,成功实现阿克曼转向机器人的仿真到现实迁移导航。

AI 中文摘要

视觉-语言导航(VLN)使机器人能够通过自然语言指令在环境中导航,使人机交互变得直观。传统的VLN模型通常依赖导航图、360度视图和完美的定位,这些在将这些模型适应现实世界环境时带来了重大挑战。本研究通过在连续环境中执行无需导航图或全景视图的VLN方法的仿真到现实域迁移,解决了这些局限性。所提出的系统集成了视觉-语言模型,将视觉输入和语言指令在共享嵌入空间中对齐,促进自然语言驱动的导航。我们采用基于跨模态注意力(CMA)的架构,在模拟环境中使用现有数据集进行训练,并使用从配备摄像头和激光雷达传感器的定制阿克曼转向机器人收集的真实世界数据对其进行微调。通过利用线性光度调整和在有限数量的片段上进行微调,我们的模型成功适应了现实世界环境,在专用硬件上离线运行的同时实现了有效的导航。使用成功率加权路径长度(SPL)和归一化动态时间规整(nDTW)指标评估的实验结果证明了我们方法的鲁棒性和适应性。关键词:视觉-语言导航,跨模态注意力,自然语言指令,仿真到现实迁移,自主导航,阿克曼转向。

英文摘要

Vision-Language Navigation (VLN) enables robots to navigate through environments using natural language instructions, making human-robot interaction intuitive. Traditional VLN models often rely on navigation graphs, 360-degree views, and perfect localization which pose significant challenges when adapting these models to real-world settings. This work addresses these limitations by performing a simulation-to-real domain shift of a VLN approach that operates in continuous environments without requiring navigation graphs or panoramic views. The proposed system integrates vision-language models that align visual inputs and linguistic instructions within a shared embedding space, facilitating natural language-driven navigation. We employ a Cross-Modal Attention (CMA) based architecture trained on an existing dataset in a simulated environment and fine-tune it using real-world data collected from a custom-built Ackermann-steered robot equipped with a camera and a LiDAR sensor. By utilising linear photometric adjustments and fine-tuning on a limited number of episodes, our model successfully adapts to real-world environments, achieving effective navigation while running offline on dedicated hardware. Experimental results, evaluated using Success weighted by Path Length (SPL) and Normalized Dynamic Time Warping (nDTW) metrics, demonstrate the robustness and adaptability of our approach. Keywords: Vision-Language Navigation, Cross-Modal Attention, Natural Language Instructions, Sim-to-Real Transfer, Autonomous Navigation, Ackermann-steering.

DOI:10.1109/ICCAR69571.2026.11549553

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑