arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11818cs.CVcs.AI

MM-ToolSandBox:一个用于评估视觉工具调用智能体的统一框架

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

Kaixin Ma, Di Feng, Alexander Metz, Jiarui Lu, Eshan Verma, Afshin Dehghan

首次发表
浏览论文内容

中文总结 AI 辅助

介绍MM-ToolSandBox框架,它能支持多图像、多轮任务,通过自动化管道生成场景。评估12个先进模型发现当前模型视觉工具调用能力不足,视觉精度是瓶颈,不同规模模型失败原因不同,该框架和基准已公开。

中文摘要 AI 辅助

我们介绍了MM-ToolSandBox,这是一个用于视觉基础工具调用智能体的基准测试和评估框架。该框架提供了一个有状态的执行环境,涵盖16个应用领域的500多个工具,支持多图像、多轮任务,智能体必须将逐渐到来的视觉输入转化为可执行的工具调用,同时处理现实的对话现象(目标修订、纠错、状态突变)。一个自动化场景生成管道通过信息流引导的规划和多阶段质量过滤产生多样化的、视觉基础的场景,产生258个人工验证的标称场景和50个针对交互式UI应用的变体。对12个从4B开放权重到前沿专有系统的先进模型进行评估表明,当前模型仍缺乏强大的视觉工具调用能力:即使是最好的模型成功率也低于50%。我们的失败分析进一步表明,视觉精度不仅是规划,也是有能力的模型的主要瓶颈:53%的失败源于从图像中提取信息不正确,尽管任务工作流程其他方面正确。随着规模的扩大,出现了规划到精度的交叉:较小的模型在决定做什么时失败,而较大的模型在感知它们所看到的东西时失败,这表明在不同能力水平上改进模型有根本不同的研究方向。该框架和基准可在这个https网址公开获得

英文摘要

We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state mutations). An automated scenario generation pipeline produces diverse, visually grounded scenarios through information-flow-guided planning and multi-stage quality filtering, yielding 258 human-verified nominal scenarios and 50 variants targeting interactive UI applications. Evaluating 12 state-of-the-art models, from 4B open-weight to frontier proprietary systems, shows that current models still lack robust visual tool-calling capability: even the best model achieves below 50% success rate. Our failure analysis further reveals that visual precision, not only planning, is a primary bottleneck for capable models: 53% of failures stem from incorrect information extraction from images despite otherwise correct task workflows. A planning-to-precision crossover emerges with scale: smaller models fail at deciding what to do, while larger models fail at perceiving what they see, suggesting fundamentally different research directions for improving models at different capability levels. The framework and the benchmark are publicly available at https://github.com/apple/ml-mmtoolsandbox

发表机构

  • Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑