arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AeroGround:用于空-地协同推理的综合基准

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

Shenghong Yi, Lin Zhang, Muzian Li, Jiakang Yuan, Haoyu Zhang, Peng Ye, Jiayuan Fan, Huafeng Qin, Tao Chen

arXiv 2608.14721首次发表:更新:

发表机构

Shanghai Innovation Institute; College of Future Information Technology, Fudan University; College of Intelligent Robotics and Advanced Manufacturing, Fudan University; Chongqing Technology and Business University; The Chinese University of Hong Kong(上海创新研究院; 复旦大学未来信息技术学院; 复旦大学智能机器人与先进制造学院; 重庆工商大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AeroGround基准,针对现有VLMs在空-地协同场景推理能力的研究空白,构建含29000组多模态观测与2250个问答实例的数据集,实验发现模型与人类表现差距显著,为相关系统开发提供基础。

AI 中文摘要

视觉语言模型(VLMs)已被广泛应用于无人机(UAV)的理解与推理任务中。现有无人机基准主要聚焦于空视图场景,但当前VLMs在空-地协同场景(如救援、基础设施巡检等实际应用场景)的理解与推理任务中能否表现良好,仍未得到充分探索。为填补这一空白,本文提出AeroGround,一个用于评估VLMs空-地协同推理能力的综合基准。AeroGround构建于模拟空-地数据集之上,包含来自多样开放环境的约29000个多模态观测组,并提供2250个高质量问答实例,涵盖跨视图对应、空间理解与推理任务。对16个预训练VLMs及两个领域适配变体的实验显示,当前模型与人类表现存在显著差距:最优模型平均准确率为54.4%,而人类达到93.3%。通过系统揭示现有模型在空-地协同推理中的优势与局限,AeroGround为开发更强大的空-地协同具身智能系统提供了基础。

英文摘要

Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmarks primarily focus on aerial-view scenarios. However, whether current VLMs can perform well on understanding and reasoning tasks in aerial-ground collaborative scenarios which are practical in real-world applications like rescue and infrastructure inspection remains underexplored. To address this gap, we introduce AeroGround, a comprehensive benchmark for evaluating VLMs in aerial-ground collaborative reasoning. AeroGround is built upon a simulated aerial-ground dataset containing approximately 29,000 multimodal observation groups from diverse open environments, and provides 2,250 high-quality question-answering instances covering cross-view correspondence, spatial understanding, and reasoning. Experiments on 16 pretrained VLMs, together with two domain-adapted variants, reveal a substantial gap between current models and human performance: the best model achieves an average accuracy of 54.4%, whereas humans reach 93.3%. By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑