arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向具身空地协同目标搜索:基准、数据集与智能体方法

Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method

Boao Yu, Zimo Chen, Junreng Rao, Yue Hu, Zhengqiu Zhu, Yong Zhao, Rusheng Ju

arXiv 2609.08402首次发表:更新:

发表机构

National University of Defense Technology(国防科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出空地目标搜索任务,构建首个基准AGOS-Bench和数据集AGOS-Dataset,并设计无需训练的智能体方法AGOS-Agent,通过搜索-交接-验证协议提升多VLM的成功率并减少决策步骤。

AI 中文摘要

城市环境中的空地目标搜索(AGOS)是一项具有挑战性的具身任务,要求一架无人机(UAV)和一辆无人地面车(UGV)从多视角视觉参考中联合搜索并验证指定的目标车辆。为研究这一未被充分探索的问题,我们引入了AGOS-Bench,这是首个专门用于评估通用视觉语言模型(VLM)能否通过无人机-地面车协同整合空中发现与地面验证的基准。我们进一步提供AGOS-Dataset,作为配套资源,包含由自动流水线构建的示例轨迹。该数据集包含7.7k个情节,用于搜索不同类别和属性的目标,涵盖三个难度级别。为解决AGOS任务,我们提出了AGOS-Agent,一种无需训练且工具增强的方法。该智能体方法通过精心设计的搜索-交接-验证协作协议,使VLM摆脱复杂动态协调的负担,仅要求VLM进行场景理解和决策。在九个VLM上的大量实验表明,AGOS-Agent在九个评估骨干网络中的八个上提高了整体成功率,同时减少了所有九个骨干网络的决策步骤。在困难子集上,Gemini-3.6-Flash的SR和SPL分别从8.6%提高到55.7%,从7.6%提高到44.0%。

英文摘要

Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV) to jointly search for and verify a specified target vehicle from multi-view visual references. To study this underexplored problem, we introduce AGOS-Bench, the first dedicated benchmark for evaluating whether general-purpose Vision-Language Models (VLMs) can integrate aerial discoveries and ground-level verification through UAV-UGV cooperation. We further provide AGOS-Dataset as the companion resource of exemplary trajectories constructed by an automatic pipeline. It consists of 7.7k episodes for searching objects of diverse categories and attributes, spanning three difficulty levels. To address the AGOS task, we propose AGOS-Agent, a training-free and tool-augmented approach. The agentic method relieves VLMs from complex and dynamic coordination via a deliberate search-handoff-verify cooperation protocol, only demanding VLMs for scene understanding and decision-making. Extensive experiments on nine VLMs show that AGOS-Agent improves overall success rate for eight of the nine evaluated backbones while reducing decision steps for all nine. On the hard split, the SR and SPL of Gemini-3.6-Flash increase from 8.6% to 55.7% and from 7.6% to 44.0%, respectively.

Comments16 pages, 4 figures, 4 tables; includes an appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑