arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08281cs.AIcs.RO

探索大语言模型(LLM)在现实世界海事航行场景中的情境理解与《国际海上避碰规则》(COLREG)遵从能力

Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios

Julius Wirbel, P. Nicholas Hansen, Line K. H. Clemmensen, Roberto Galeazzi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探索LLM在海事航行中的应用,构建含50个真实场景的AIS数据集,评估不同LLM的理解与推理能力,发现不微调则难以解决海事航行任务。

中文摘要 AI 辅助

近年来,大语言模型(LLM)在不同领域展现出强大的情境理解、推理与决策能力,在汽车领域尤为突出。因此,本研究探索当前最先进的LLM作为海事航行工具,涵盖《国际海上避碰规则》(COLREGs)等成文规则,以及“良好航海技艺”概念所总结的不成文最佳实践。我们构建了一个包含50种来自自动识别系统(AIS)数据的多样化现实航行场景的数据集,为场景标注适用的COLREG规则、推荐行动及行动推理。我们探索了多种不同的LLM架构与规模,以确定它们对海事航行任务的理解能力,并评估其在该领域的推理能力。所得结果表明,即便是对于较大的在线模型,若不进行微调,海事航行任务仍难以解决。

英文摘要

Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in different domains, most notable in the automotive sector. Therefore, we explore current state-of-the-art LLMs as a tool for maritime navigation, which includes both codified rules in the Collision Regulations (COLREGs) and uncodified best practices summarized in the concept of ``Good Seamanship''. We construct a dataset consisting of 50 diverse, real-world navigation scenarios from AIS data, label scenarios with applicable COLREG rules, recommended actions, and the reasoning for the action. We explore a variety of different LLM architectures and sizes to determine their understanding of maritime navigation tasks as well as evaluate their reasoning capabilities in this domain. The results obtained indicate that the maritime navigation task remains difficult to solve without fine-tuning, even for larger online models.

发表机构

  • Technical University of Denmark(丹麦技术大学)
  • University of Copenhagen(哥本哈根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑