评估多模态大语言模型(MLLM)作为通用视觉-语言-动作智能体在无人机控制中的应用:指挥、接近、跟踪与搜索
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
浏览论文内容
中文总结 AI 辅助
本文提出DroneCATS-Agent架构与DroneCATS基准,评估不同规模MLLM作为通用视觉-语言-动作智能体在无人机控制中的四项核心能力,发现模型动作协议缺陷是关键差距,为缩小该差距提供了评估基准。
中文摘要 AI 辅助
多模态大语言模型(Multimodal Large Language Models, MLLM)是强大的图像与视频感知模型,本文探究其在动作领域的应用深度:将MLLM直接接入无人机控制环路,仅通过提示词声明其全部动作空间。现有相关系统虽已接近该场景,但逐渐限制模型的决策能力,本文则将其决策范围拓宽。我们提出DroneCATS-Agent架构,其中MLLM为可替换组件;同时构建DroneCATS基准,将模型作为自变量进行评估。该智能体不仅能飞向像素点,还可控制无人机偏航、执行搜索、在不确定时弃权(不执行)、自行声明到达,且无需微调或函数调用模式。我们评估前沿模型与开源模型在四项核心能力上的表现:接近可见目标、跟踪移动目标、在初始视野外搜索、指挥多无人机编队。结果显示,即使是最简单的具身场景也远未解决;关键是,为识别边缘场景中最先失效的环节,我们将模型规模缩小至2B参数。研究发现揭示了一个显著悖论:失效的并非飞行环节——小型开源模型通常比前沿模型更可靠地进入成功半径,却因过早或未声明到达而导致任务失败;多无人机指挥会放大这种差距,小型模型会因在不同视角间盲目复制单一坐标而失效。作为视觉-语言-动作智能体,模型的空间感知能力尚可,但动作协议存在缺陷。可部署的边缘模型与前沿模型的区别不在于导航,而在于维持声明协议并发出正确终止动作的能力。尚未解决的问题是在机载计算成本下缩小这一差距,以得到一个能持续规划且明确何时完成的快速模型,而DroneCATS基准正是为衡量这一差距而构建。
英文摘要
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control loop, with its entire action space declared solely in the prompt. Recent systems approach this setting but increasingly narrow the model's decision-making. We widen it back. We introduce DroneCATS-Agent, an architecture where the MLLM is a swappable component, and DroneCATS, a benchmark treating the model as the independent variable. Beyond merely flying toward a pixel, our agent entrusts the model to yaw and search, deliberate when unsure, and self-declare arrival---all without fine-tuning or function-calling schemas. Evaluating frontier and open models across four core capabilities---approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet---reveals that even the simplest embodied settings are far from solved. Crucially, to identify what breaks first at the edge, our roster scales down to 2B parameters. The findings expose a stark paradox: it is not the flying that fails. Small open models often navigate into the success radius more reliably than frontier models, yet lose the episode by declaring arrival prematurely or not at all. Multi-drone commanding amplifies this divide, with small models failing by blindly copying a single coordinate across distinct views. Viewed as vision-language-action agents, the models' spatial perception holds up, but their action protocol does not. What separates a deployable edge model from a frontier model is not navigation, but the discipline to sustain a declared protocol and emit the correct terminating action. The open problem is closing this gap at onboard compute costs---yielding a fast model that plans persistently and knows exactly when it is done---and DroneCATS is built to measure that distance.
发表机构
- NAVER Cloud(NAVER云)
机构由 AI 辅助整理,请以论文原文为准。