arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11738cs.CVcs.AI

推进基于多模态大语言模型(MLLM)的无人机(UAV)图像理解与推理:基准测试及无训练多智能体系统

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

Haoyu Zhang, Shuoxun Zhang, Peng Ye, Lin Zhang, Jiakang Yuan, Shenghong Yi, Yuening Wang, Tao Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建UAVQA-Bench基准,识别MLLM用于无人机图像理解的三类失效模式,提出含DSPE、CAIR、DAAS的无训练多智能体系统UAV-MAS,其32B版本在基准上准确率超Gemini 3 Pro 4.0个百分点

中文摘要 AI 辅助

基于多模态大语言模型(MLLM)的无人机(UAV)航拍图像理解与推理是空天智能的核心需求,但面临极端尺度变化、相机姿态任意、目标密度高等独特挑战。尽管相关研究日益增多,现有评估仍分散在不同数据集和狭窄任务中,导致缺乏对无人机理解与推理能力的统一评估,存在关键缺口。为填补该缺口,本研究构建了UAVQA-Bench基准,该基准包含1500个人工标注的问答对,源自13个公开无人机数据集,覆盖6个能力维度及16项任务,涵盖选择题与视觉定位两种格式。通过对广泛的开源、闭源MLLM及基于智能体的系统在UAVQA-Bench上的系统评估,识别出三类关键失效模式:领域工具集不匹配、未受管控的错误传播、静态推理。基于这些发现,本研究提出UAV-MAS,一种用于基于MLLM的无人机航拍图像理解与推理的无训练多智能体系统,包含三个核心组件:领域专用感知引擎(DSPE),可将查询路由至任务适配的视觉工具;上下文感知迭代优化模块(CAIR),可验证中间推理以抑制错误累积;难度感知自适应搜索机制(DAAS),可根据问题难度调整搜索深度。采用32B参数开源MLLM的UAV-MAS在UAVQA-Bench上的总体准确率达77.0%,超过Gemini 3 Pro 4.0个百分点;而8B参数版本较其基础模型提升8.7个百分点。

英文摘要

Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme scale variation, arbitrary camera orientations, and high object density. Despite growing interest, existing evaluations remain fragmented across individual datasets and narrow tasks, leaving a critical gap in unified assessment of UAV understanding and reasoning capabilities. To fill this gap, we construct UAVQA-Bench, a benchmark of 1,500 human-annotated QA pairs drawn from 13 public UAV datasets, covering 6 capability dimensions and 16 tasks in both multiple-choice and visual grounding formats. Systematic evaluation of a broad range of open-source and closed-source MLLMs as well as agent-based systems on UAVQA-Bench identifies three key failure modes: domain-toolset mismatch, unchecked error propagation, and static reasoning. Motivated by these findings, we propose UAV-MAS, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine (DSPE) that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error accumulation, and a Difficulty-Aware Adaptive Search mechanism (DAAS) that adjusts search depth to question difficulty. UAV-MAS with a 32B open-source MLLM achieves 77.0% overall accuracy on UAVQA-Bench, surpassing Gemini 3 Pro by 4.0\%, while the 8B variant improves 8.7\% over its base model.

发表机构

  • Fudan University(复旦大学)
  • College of Future Information Technology(未来信息技术学院)
  • Shanghai Innovation Institute(上海创新研究院)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑