arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AerialDojo-200K:面向开放世界空中目标搜索的大规模基准套件

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu

arXiv 2609.36066首次发表:更新:

发表机构

Tsinghua University; University College Dublin; Peking University; Zhejiang University of Technology(清华大学; 都柏林大学; 北京大学; 浙江工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对开放世界空中目标搜索缺乏大规模基准的问题,提出AerialDojo-200K,包含42个场景和20余万任务实例,并构建统一评估框架,揭示现有模型与通用智能体的差距。

AI 中文摘要

开放世界空中目标搜索是一项基础但具有挑战性的任务,要求空中智能体自主探索大规模、非结构化的三维环境,并根据语义描述或参考图像到达目标物体,而非遵循特定路线的指令。然而,该任务的研究仍处于起步阶段,依赖于规模较小、针对特定环境的基准,且这些基准具有异构的动作空间和数据格式。这些局限性阻碍了大规模训练和跨基准评估,限制了空中智能体的可扩展性和泛化能力。为解决这一问题,我们提出了AerialDojo-200K,一个面向开放世界空中目标搜索的大规模基准套件,其场景数量是现有最大基准的3倍,任务实例数量是其18.7倍。具体而言,我们构建了42个仿真场景,涵盖四个场景族和21种场景类型,包括18个城市场景、12个自然场景、6个基础设施场景和6个灾难场景。为确保数据质量,12名标注员花费两个月时间在这些场景中手动标注了109个地标、2099个目标物体和2099个物体锚点。我们进一步构建了205,732个任务实例,包括超过10万个语义目标实例和超过10万个图像目标实例,涵盖基础、标准和长时程设置。每个任务实例包含一条无碰撞参考轨迹和相应的多视角视频记录。我们还开发了一个统一的评估框架,场景划分包括21个分布内场景和21个分布外场景。最后,我们对五个开源和四个闭源多模态大语言模型的评估表明,实现通用空中智能体仍有很长的路要走。所有内容可在以下网址找到:https://this https URL。

英文摘要

Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes. We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents. All can be found at https://fengtt42.github.io/AerialDojo/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑