arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2508.13142cs.CVcs.CLcs.LGcs.MMcs.RO

多模态大语言模型在空间智能方面的综合评估

Holistic Evaluation of Multimodal LLMs on Spatial Intelligence

  • SenseTime Research(商汤科技研究院)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhongang Cai, Yubo Wang, Qingping Sun, Ruisi Wang, Chenyang Gu, Wanqi Yin, Zhiqian Lin, Zhitao Yang, Chen Wei, Oscar Qian, Hui En Pang, Xuanke Shi, Kewang Deng,… 展开作者

Zhongang Cai, Yubo Wang, Qingping Sun, Ruisi Wang, Chenyang Gu, Wanqi Yin, Zhiqian Lin, Zhitao Yang, Chen Wei, Oscar Qian, Hui En Pang, Xuanke Shi, Kewang Deng, Xiaoyang Han, Zukai Chen, Jiaqi Li, Xiangyu Fan, Hanming Deng, Lewei Lu, Bo Li, Ziwei Liu, Quan Wang, Dahua Lin, Lei Yang

更新

AI总结:

本文提出EASI框架,用于评估多模态大语言模型在空间智能方面的综合表现,揭示GPT-5在SI中的优势与不足,并展示专有模型在困难任务上的劣势。

AI中文摘要:

近年来,多模态模型在多方面取得了显著进展。然而,它们在空间理解和推理方面仍然存在明显的局限性,这种能力是将人工智能通用智能锚定在物理世界中的关键。随着GPT-5的最新发布,号称目前最强大的AI模型,现在有必要考察领先模型(GPT、Gemini、Grok、Seed、Qwen和Intern)在空间智能(SI)道路上的位置。为此,我们提出了EASI,用于多模态大语言模型在空间智能方面的综合评估。EASI概念化了一种全面的空间任务分类,统一了现有的基准测试并整合了越来越多的新整理的测试集,使能够系统评估最先进模型。在本报告中,我们跨八个关键基准进行研究,总成本超过十亿个token。我们的实证研究揭示了(1)GPT-5在SI方面表现出前所未有的强大,但(2)在广泛的空间任务范围内仍然显著落后于人类表现。此外,我们(3)显示SI任务暴露了比非SI任务更大的模型能力缺陷,到这种程度(4)专有模型在面对最困难的任务时并不具有决定性优势。此外,我们对一组多样化的场景进行了定性评估,这些场景对人类来说是直观的,但最先进多模态模型却无法通过。EASI是一个持续的社区努力:我们已经开源了EASI代码库,提供了一站式和可重复的解决方案,具有标准化的接口、集成的协议和提示,显著减少了配置和运行多个基准测试的摩擦;我们还推出了一个配套的EASI排行榜,提供模型在完整SI范围内的持续更新性能快照,加速集体在稳健SI方面的进步。

英文摘要:

Multimodal models have achieved remarkable progress in recent years. Nevertheless, they continue to exhibit notable limitations in spatial understanding and reasoning, the very capability that anchors artificial general intelligence in the physical world. With the recent release of GPT-5, allegedly the most powerful AI model to date, it is timely to examine where the leading models (GPT, Gemini, Grok, Seed, Qwen, and Intern) stand on the path toward spatial intelligence (SI). We thus propose EASI for holistic Evaluation of multimodAl LLMs on Spatial Intelligence. EASI conceptualizes a comprehensive taxonomy of spatial tasks that unifies existing benchmarks and a growing collection of newly curated ones, enabling systematic evaluation of state-of-the-art models. In this report, we conduct the study across eight key benchmarks, at a cost exceeding ten billion total tokens. Our empirical study then reveals that (1) GPT-5 demonstrates unprecedented strength in SI, yet (2) still falls short of human performance significantly across a broad spectrum of SI-tasks. Moreover, we (3) show that SI-tasks expose greater model capability deficiency than non-SI tasks, to the extent that (4) proprietary models do not exhibit a decisive advantage when facing the most difficult ones. In addition, we conduct a qualitative evaluation across a diverse set of scenarios that are intuitive for humans, yet fail the most advanced multimodal models. EASI is an ongoing community effort: we have open-sourced the EASI codebase that provides a one-stop and reproducible solution with standardized interfaces, integrated protocols and prompts that significantly reduce the friction of configuring and running multiple benchmarks; we have also launched an accompanying EASI leaderboard to provide a continually updated snapshot of model performance across the full SI spectrum, accelerating collective progress toward robust SI.

补充信息

↑