AI 中文总结
研究空中MLLMs在零样本下解决长期具身任务的能力,引入MissionBench基准,含120个任务。对比22个模型与人类表现,发现模型成功率低,虽规模扩大有收益,但任务级能力需协调多种能力,凸显具身AI改进的前景与风险。
AI 中文摘要
多模态大语言模型(MLLMs)正成为具身智能体的核心推理模块,但通用模型能否从单个高级指令解决长期具身任务尚不明晰。我们引入了MissionBench,这是一个用于空中3D环境中MLLMs任务级评估的基准。它包含五个模拟3D环境和四个任务家族中的120个任务。智能体必须仅使用自我中心观察及其动作历史自主规划、导航并报告结果,无需特定于空中的微调。在22个开源和闭源MLLMs中,最强模型的任务成功率不到35%,而人类表现为84.4%,凸显了多步具身任务的难度。尽管模型家族间差异很大,但我们观察到规模扩大带来的收益,表明更大的通用模型具有更强的零样本具身能力。我们的分析表明,任务级能力需要协调空间感知之外的多种能力,包括多步规划和自适应推理。这推动了闭环评估,并凸显了具身人工智能中规模驱动改进的前景和风险。
英文摘要
Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction. We introduce MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments. It comprises 120 missions across five simulated 3D environments and four task families. Agents must autonomously plan, navigate, and report outcomes using only egocentric observations and its action history, without aerial-specific fine-tuning. Across 22 open- and closed-source MLLMs, the strongest model succeeds on fewer than 35% of missions compared to 84.4% human performance, highlighting the difficulty of multi-step embodied tasks. Despite large variations between model families, we observe gains from scaling, indicating that larger general-purpose models possess stronger zero-shot embodied capabilities. Our analysis shows that mission-level competence requires coordinating multiple capabilities beyond spatial perception, including multi-step planning and adaptive reasoning. This motivates closed-loop evaluation and highlights both the promise and risk of scaling-driven improvements for embodied AI.
CommentsPreprint