Agentic-MME:代理能力为多模态智能带来了什么?
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
- CASIA(中国科学院自动化研究所)
- UCAS(中国科学院大学)
- SEU(东南大学)
- NJU(南京大学)
- PKU(北京大学)
- BUAA(北京航空航天大学)
- NTU(南洋理工大学)
- UCLA(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Agentic-MME通过418个真实任务评估多模态代理能力,采用过程验证方法,量化效率并验证工具调用准确性,实验显示模型在复杂任务中表现显著下降。
AI中文摘要:
多模态大语言模型(MLLMs)正从被动观察者发展为主动代理,通过视觉扩展(调用视觉工具)和知识扩展(开放网络搜索)解决问题。然而现有评估不足:缺乏灵活工具集成、单独测试视觉和搜索工具、主要通过最终答案评估。为此,我们引入Agentic-MME,一个用于多模态代理能力的过程验证基准。它包含6个领域、3个难度级别的418个真实任务,包含超过2000个逐步检查点,每个任务平均需10多人小时的手动标注。每个任务包含统一的评估框架,支持沙盒代码和API,以及标注了双轴(S轴和V轴)的人类参考轨迹。为实现真正的过程级验证,我们审计细粒度的中间状态而非仅最终答案,并通过相对于人类轨迹的过度思考度量效率。实验结果表明,最佳模型Gemini3-pro在整体准确率上达到56.3%,但在Level-3任务中显著下降至23.0%,凸显了真实世界多模态代理问题解决的难度。
英文摘要:
Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall short: they lack flexible tool integration, test visual and search tools separately, and evaluate primarily by final answers. Consequently, they cannot verify if tools were actually invoked, applied correctly, or used efficiently. To address this, we introduce Agentic-MME, a process-verified benchmark for Multimodal Agentic Capabilities. It contains 418 real-world tasks across 6 domains and 3 difficulty levels to evaluate capability synergy, featuring over 2,000 stepwise checkpoints that average 10+ person-hours of manual annotation per task. Each task includes a unified evaluation framework supporting sandboxed code and APIs, alongside a human reference trajectory annotated with stepwise checkpoints along dual-axis: S-axis and V-axis. To enable true process-level verification, we audit fine-grained intermediate states rather than just final answers, and quantify efficiency via an overthinking metric relative to human trajectories. Experimental results show the best model, Gemini3-pro, achieves 56.3% overall accuracy, which falls significantly to 23.0% on Level-3 tasks, underscoring the difficulty of real-world multimodal agentic problem solving.