arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26084cs.RO

视觉语言模型作为自主无人机导航的副驾驶:退化环境中的延迟与可靠性分析

Vision-Language Models as copilots for Autonomous UAV Navigation: Analysis of Latency and Reliability in Degraded Environments

发表机构乌拉圭科技大学
查看机构详情
  • Technological University of Uruguay(乌拉圭科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Hiago Sodre, Sebastian Barcelona, Vincent Sandin, Pablo Moraes, Ahilen Mazondo, Igor Nunes, William Moraes, André Kelbouscas, Ricardo Grando

首次发表
浏览论文内容

中文总结 AI 辅助

针对无人机自主导航中VLM的实时性与可靠性挑战,提出FSM-VLM混合架构,通过SITL仿真发现参数规模是集成的主要瓶颈。

中文摘要 AI 辅助

将视觉语言模型(VLM)集成到自主无人机(UAV)中,提供了前所未有的语义推理能力。然而,实时闭环导航不仅需要低推理延迟,还需要服从结构化飞行指令。本文提出了一种用于无GPS环境下无人机的混合FSM-VLM控制架构。该系统将用于低级物理控制的确定性有限状态机(FSM)与用于高级语义路径规划的异步VLM副驾驶相结合。我们在软件在环(SITL)仿真中评估了三种不同参数规模的模型。该框架在正常和退化场景中,分别隔离并测量了格式层面的语法错误与逻辑层面的语义幻觉。研究表明,参数规模而非纯延迟,仍是VLM安全且兼容地集成到自主飞行的主要瓶颈。

英文摘要

The integration of Vision-Language Models (VLMs) in autonomous Unmanned Aerial Vehicles (UAVs) offers unprecedented semantic reasoning capabilities. However, real-time closed-loop navigation requires not only low inference latency but also obedience to structured flight commands. This paper proposes a hybrid FSM-VLM control architecture for UAVs in GPS-free environments. The system combines a deterministic Finite State Machine (FSM) for low-level physical control with an asynchronous VLM copilot for high-level semantic pathfinding. We evaluate three models with different parameter scales in a Software-In-The-Loop (SITL) simulation. The framework isolates and measures syntax errors at the format level versus semantic hallucinations at the logic level in a normal and degraded scenario. This study demonstrates that parameter scaling, and not pure latency, remains the primary bottleneck for the safe and compatible integration of VLM into autonomous flights.

↑