视觉语言模型作为自主无人机导航的副驾驶:退化环境中的延迟与可靠性分析
Vision-Language Models as copilots for Autonomous UAV Navigation: Analysis of Latency and Reliability in Degraded Environments
查看机构详情
- Technological University of Uruguay(乌拉圭科技大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对无人机自主导航中VLM的实时性与可靠性挑战,提出FSM-VLM混合架构,通过SITL仿真发现参数规模是集成的主要瓶颈。
中文摘要 AI 辅助
将视觉语言模型(VLM)集成到自主无人机(UAV)中,提供了前所未有的语义推理能力。然而,实时闭环导航不仅需要低推理延迟,还需要服从结构化飞行指令。本文提出了一种用于无GPS环境下无人机的混合FSM-VLM控制架构。该系统将用于低级物理控制的确定性有限状态机(FSM)与用于高级语义路径规划的异步VLM副驾驶相结合。我们在软件在环(SITL)仿真中评估了三种不同参数规模的模型。该框架在正常和退化场景中,分别隔离并测量了格式层面的语法错误与逻辑层面的语义幻觉。研究表明,参数规模而非纯延迟,仍是VLM安全且兼容地集成到自主飞行的主要瓶颈。
英文摘要
The integration of Vision-Language Models (VLMs) in autonomous Unmanned Aerial Vehicles (UAVs) offers unprecedented semantic reasoning capabilities. However, real-time closed-loop navigation requires not only low inference latency but also obedience to structured flight commands. This paper proposes a hybrid FSM-VLM control architecture for UAVs in GPS-free environments. The system combines a deterministic Finite State Machine (FSM) for low-level physical control with an asynchronous VLM copilot for high-level semantic pathfinding. We evaluate three models with different parameter scales in a Software-In-The-Loop (SITL) simulation. The framework isolates and measures syntax errors at the format level versus semantic hallucinations at the logic level in a normal and degraded scenario. This study demonstrates that parameter scaling, and not pure latency, remains the primary bottleneck for the safe and compatible integration of VLM into autonomous flights.