arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向近海风电场区域自主水下航行器(AUV)与自主水面航行器(ASV)导航的世界模型赋能大语言模型规划

World-Model-Grounded LLM Planning for AUV and ASV Navigation Near Offshore Wind Farms

Markus Buchholz, Ignacio Carlucho, Yvan R. Petillot

arXiv 2608.19661首次发表:更新:

发表机构

Norwegian Defence Research Establishment (FFI); School of Engineering & Physical Sciences, Heriot-Watt University(挪威国防研究机构(FFI); 赫瑞瓦特大学工程与物理科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出世界模型赋能大语言模型规划方法,结合MPC闭环重规划器等组件,在近海风电场区域的AUV和ASV导航任务中,大幅降低目标距离误差,验证了其有效性。

AI 中文摘要

大语言模型可将自然语言任务转化为机器人动作序列,但缺乏物理感知能力,无法判断指令的执行时长,也无法判断是否会使机器人漂移至障碍物处。我们提出利用世界模型扩展基于大语言模型的规划器的能力,该方法包含三个组件:基于物理的神经世界模型、三阶段梯度轨迹优化器,以及带有信赖域保护的模型预测控制器(MPC)式闭环重规划器。语言模型决定执行什么操作,世界模型决定执行时长,无论是驱动6自由度(DOF)的8个推进器,还是3自由度的两个差速推进器。我们对两种在近海风电场基础设施附近作业的海洋航行器进行评估:6自由度自主水下航行器(AUV)和3自由度差速驱动自主水面航行器(ASV)。在每种航行器的5项基准任务中,两种航行器均达成所有目标且无预测碰撞;在经过残差微调后,替代展开的均方根误差(RMSE)分别降低60%(AUV)和69%(ASV),随后将模型迁移至GazeboSim仿真环境,在洋流、波浪和推进器动力学条件下仍保持无碰撞,且与未赋能基线相比,GazeboSim的目标距离误差降低70%-82%(ASV)和约93%(AUV)。针对ASV,我们还展示了视觉语言模型(VLM)辅助的语义建图流水线,该流水线从卫星图像、航海图和预测应用程序编程接口(API)中提取障碍物和环境上下文,而非依赖机载传感器,作为手动指定障碍物几何结构的替代方案,其导航准确率达96%。

英文摘要

Large language models can turn a natural-language mission into a sequence of robot actions, but they do not have a sense of physics: they cannot judge how long a command should run, or whether it will make the robot drift into an obstacle. We proposed the use of a world model to expand the capabilities of Large Language model-based planners. Our method has three components: a physics-grounded neural world model, a three-phase gradient-based trajectory optimizer, and a Model Predictive Controller (MPC)-style closed-loop replanner with a trust-region guard. The language model decides what to do, and the world model decides how long, whether that means driving eight thrusters through 6 DOF or two differential thrusters through 3 DOF. We evaluate two marine vehicle classes operating near offshore wind infrastructure: a 6-DOF Autonomous Underwater Vehicle (AUV) and a 3-DOF differential-drive Autonomous Surface Vehicle (ASV). In five benchmark missions per platform, both vehicles reach every goal with zero predicted collisions, and both transfer to GazeboSim under ocean current, waves, and thruster dynamics, remaining collision-free and cutting GazeboSim goal-distance error versus the ungrounded baseline by 70-82% (ASV) and roughly 93% (AUV), after a residual fine-tuning pass that separately reduces surrogate rollout Root Mean Square Error (RMSE) by 60% (AUV) and 69% (ASV). For the ASV we further demonstrate a Vision language model (VLM)-assisted semantic-mapping pipeline that extracts obstacles and environmental context from satellite imagery, nautical charts, and forecast Application Programming Interface (API) instead of onboard sensors, reaching 96% navigability accuracy as a drop-in replacement for hand-specified obstacle geometry.

CommentsThis work has been accepted to the IEEE IROS 2026 AQ2UASIM workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑