arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mine Odyssey:野外智能体空间智能基准测试

Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild

Yuxuan Cao, Junlong Li, Hao Li, Junxian He

arXiv 2610.11328首次发表:更新:

发表机构

The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出基于Minecraft真实地点重建的智能体空间智能基准Mine Odyssey,含180个任务,评估显示GPT-6 Astra成功率85.6%,开源模型DeepSeek-V4.1-Flash仅23.9%,为该领域发展提供了研究基础。

AI 中文摘要

基础模型的进展正推动将智能体引入现实世界以辅助人类,这类智能体需要具备智能体空间智能:探索陌生环境、通过交互更新空间理解、基于反馈调整行动以持续达成一系列目标。现有基准仅覆盖有限的空间布局、规模和遍历需求。我们推出Mine Odyssey,一个利用Minecraft重建的真实世界地点来评估智能体空间智能的基准,包含180个任务,覆盖五大洲20个国家和地区的30个地点,其中包括20个户外场景和10个室内场景。这些场景涵盖多样的空间规模、布局、地形和连通模式,从中城曼哈顿、农村恩特鲁普到圣卢西亚山、白金汉宫。我们选择有意义的途经点,如地标、建筑和房间,并手动验证其可达性。每个任务提供自然语言指令,指定需访问的途经点及其顺序。完成这些任务需要智能体找到可达路线和入口、开门、通过楼梯和梯子在不同楼层间移动,同时监控自身进度并从导航错误中恢复。在评估的8个最先进模型中,GPT-6 Astra达到最高成功率85.6%;排名第二的模型Claude Opus 5.5完成73.9%的任务;最强的开源权重模型DeepSeek-V4.1-Flash达到23.9%,凸显当前模型在智能体空间智能方面仍有巨大提升空间。对Mine Odyssey的综合分析和消融研究揭示了当前模型的局限性,并为推进智能体空间智能提供了见解。

英文摘要

Advances in foundation models are driving efforts to introduce agents to assist people in the physical world. Such agents require agentic spatial intelligence: exploring unfamiliar environments, updating spatial understanding through interaction, and adapting actions based on feedback to sustain progress toward a sequence of goals. Existing benchmarks cover only a limited range of spatial layouts, scales, and traversal requirements. We introduce Mine Odyssey, a benchmark for evaluating agentic spatial intelligence using Minecraft reconstructions of real-world locations. It comprises 180 tasks covering 30 such locations across 20 countries and regions on five continents, including 20 outdoor and 10 indoor settings. These settings span diverse spatial scales, layouts, terrains, and connectivity patterns, from Midtown Manhattan and rural Entrup to Santa Lucía Hill and Buckingham Palace. We select meaningful waypoints, such as landmarks, buildings, and rooms, and manually verify their accessibility. Each task provides a natural-language instruction specifying which waypoints to visit and in what order. Completing these tasks requires agents to find accessible routes and entrances, open doors, and move between levels using stairs and ladders, while monitoring their progress and recovering from navigation errors. Across eight evaluated state-of-the-art models, GPT-6 Astra achieves the highest success rate of 85.6%. However, the second-best model, Claude Opus 5.5, completes 73.9% of tasks, while the strongest evaluated open-weight model, DeepSeek-V4.1-Flash, reaches 23.9%, highlighting substantial room for improvement in the agentic spatial intelligence of current models. Comprehensive analyses and ablation studies on Mine Odyssey reveal current models' limitations and provide insights for advancing agentic spatial intelligence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑