arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20620cs.ROcs.AI

AUV故障恢复仿真平台:探索基于LLM的诊断策略

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

Khalid Halba, Kylie Cooper, James G. Bellingham

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SPAR仿真平台,集成LLM诊断与恢复规划,通过480次试验评估不同模型,发现前沿模型诊断性能更优,贡献了从检测到缓解的恢复架构及评估方法。

中文摘要 AI 辅助

在超出可靠通信范围作业的自主水下机器人(AUV)必须无需人工干预即可从故障中恢复。我们研究了一种架构,其中传统的确定性分层控制自主性负责正常操作,而当机载异常检测识别出性能超出预期限制时,可调用的大语言模型(LLM)作为诊断和恢复规划器。由于语言模型具有随机性,严格的评估需要集成测试而非单个演示。我们提出了一种闭环仿真架构,将实时C车辆软件与用于基于物理的故障注入、结构化提示、语言模型交互、任务文件生成、验证、执行和LLM评判打分的更高级编排层相结合。该框架称为SPAR(AUV恢复仿真平台),支持跨故障实现、提示结构、推理模型和任务条件的评估。我们针对质量偏移故障在480次SPAR试验中变化这些因素,评估了一个前沿模型和三个现成的本地可部署LLM。模型选择主导诊断:前沿模型在85-90%的试验中将其重心偏移机制置于前三假设中,而最佳本地模型为60-78%。推理分析表明,本地模型的成功与遵循完整诊断程序相关,而较弱的模型往往过早地提交到升降舵故障,即使执行器跟踪其命令。在此数据集中,诊断和操作决策性能似乎不耦合。贡献在于一种将意外故障恢复从检测扩展到缓解的架构,以及一种用于在低功耗AUV上评估LLM辅助任务管理的集成方法。

英文摘要

Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnostic and recovery planner when onboard anomaly detection identifies performance outside expected limits. Because language models are stochastic, rigorous evaluation requires ensemble testing rather than individual demonstrations. We present a closed-loop simulation architecture that couples real-time C vehicle software with a higher-level orchestration layer for physics-based fault injection, structured prompting, language-model interaction, mission file generation, validation, execution, and LLM-judge scoring. The framework, which we call SPAR (Simulation Platform for AUV Recovery), supports evaluation across fault realizations, prompt structures, reasoning models, and mission conditions. We vary these for a mass-shift fault over 480 SPAR trials, evaluating a frontier model and three off-the-shelf locally deployable LLMs. Model choice dominates diagnosis: the frontier model places the CG-shift mechanism in its top three hypotheses in 85-90% of trials, versus 60-78% for the best local model. Reasoning analysis indicates that local-model success is associated with following the complete diagnostic procedure, whereas weaker models often commit prematurely to elevator failure even though the actuator tracks its command. Diagnosis and operational decision performance do not appear to be coupled in this dataset. The contributions are an architecture extending unanticipated-fault recovery from detection to mitigation and an ensemble methodology for evaluating LLM-assisted mission management on low-power AUVs.

发表机构

  • Johns Hopkins Institute for Assured Autonomy(约翰霍普金斯大学保障自主研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑