arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

认识自我,理解世界:用于基于多模态大语言模型的无人机时空推理的双认知基准测试

Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs

Like Liu, Zhengzheng Xu, Haitao He, Hongzhe Li, Shuchang Zhang, Dian Shao

arXiv 2607.16193首次发表:更新:

发表机构

Northwestern Polytechnical University; China University of Petroleum(西北工业大学; 中国石油大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多模态大语言模型在无人机场景双认知能力评估不足的问题,提出UAV-DualCog基准测试,涵盖图像和视频任务,通过自动化管道构建数据,评估发现现有模型存在瓶颈,该基准测试有价值。

AI 中文摘要

多模态大语言模型在各种视觉语言任务中取得了强大性能,但在无人机场景中的能力仍未得到充分探索。近期面向无人机的基准测试开始评估多模态大语言模型在空中场景中的表现,但通常侧重于场景理解、事件识别或导航完成,而非联合评估无人机智能体所需的双认知能力,即在多视图时空背景下对无人机自身状态和外部环境进行推理。为填补这一空白,我们提出了UAV-DualCog,这是一个基于双认知视角构建的空中多视图时空推理基准测试。UAV-DualCog包括图像和视频任务,以联合评估自我状态和环境状态推理,同时要求超越离散答案预测的空间或时间基础。我们还开发了一个自动化管道,从场景级语义点云构建数据,产生一个具有多样场景、数百个地标和数千个问答样本且可扩展的基准测试。广泛评估表明,当前多模态大语言模型在无人机双认知方面仍远不可靠。自我状态推理、视角转换、精确空间基础和时间间隔定位是持续存在的瓶颈,通过思维/前沿模型和人类基线进行的额外验证证实,该基准测试对人类来说是可理解的,但对现有模型具有挑战性。我们进一步从不相交的场景构建了UAV-DualCog-Train,并通过一个轻量级优化探针表明它提供了有用的结构化监督,这表明它不仅作为评估基准有价值,而且作为推进基于多模态大语言模型的无人机智能体的数据资源也有价值。项目网站和补充材料:此https URL

英文摘要

Multimodal large language models have achieved strong performance across diverse vision-language tasks, yet their capabilities in UAV scenarios remain insufficiently explored. Recent UAV-oriented benchmarks have begun to evaluate MLLMs in aerial scenarios, but they typically focus on scene understanding, event recognition, or navigation completion, rather than jointly assessing the dual-cognition capability required for UAV agents: reasoning about both the UAV's own state and the external environment in multiview spatio-temporal contexts. To address this gap, we present UAV-DualCog, a benchmark for aerial multiview spatio-temporal reasoning built on this dual-cognition perspective. UAV-DualCog includes both image and video tasks to jointly evaluate self-state and environment-state reasoning, while requiring spatial or temporal grounding beyond discrete answer prediction. We also develop an automated pipeline that constructs data from scene-level semantic point clouds, yielding a scalable benchmark with diverse scenes, hundreds of landmarks, and thousands of QA samples. Extensive evaluations show that current MLLMs remain far from reliable in UAV dual cognition. Self-state reasoning, viewpoint transformation, precise spatial grounding, and temporal interval localization are persistent bottlenecks, and additional validation with thinking/frontier models and a human baseline confirms that the benchmark is understandable to humans but challenging for existing models. We further construct UAV-DualCog-Train from disjoint scenes and show through a lightweight optimization probe that it provides useful structured supervision, suggesting its value not only as an evaluation benchmark but also as a data resource for advancing MLLM-based UAV agents. Project website and supplementary materials: https://uav-dualcog.lozumi.com

Comments11 pages, 4 figures, 7 tables. Project website: https://uav-dualcog.lozumi.com

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑