发表机构
Peking University; Tsinghua University; SenseTime Research(北京大学; 清华大学; 商汤科技研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出名为EASEL的基准测试,结合44万样本的EASEL-Data数据集与EASEL-9B模型,评估多模态智能体的灵巧视觉工具使用能力,发现现有25个模型在该任务上表现不佳,EASEL-9B相对基础模型提升6.3%。
AI 中文摘要
评估正从静态问答转向智能体设置,其中模型通过外部工具执行动作。我们在此领域中识别出一种关键但未被充分探索的能力——灵巧视觉工具使用:这是一种细粒度、闭环、参数化的视觉动作,模型从视觉证据中推断工具参数,且这些参数直接决定最终结果。现有基准涵盖网页导航、图形用户界面(GUI)操作和软件工程,但很少关注视觉证据与执行精度之间的耦合关系。我们提出EASEL,一个评估灵巧视觉工具使用受控实例的基准,采用参考引导的视觉重构作为主要代理任务:智能体逐步绘制画布以匹配参考图像。EASEL还包含区域标注、手写和路径规划等语义任务。我们进一步提供EASEL-Data,这是一个包含44万样本的两阶段课程数据集,用于轨迹监督,以及EASEL-9B,以研究其对该能力的影响。对25个模型的评估显示,当前多模态智能体在EASEL上普遍表现不佳:重构相似度在低水平达到瓶颈(0.40-0.54),而轨迹诊断暴露出严重的闭环不稳定性——模型通常会过早饱和或在峰值后性能下降;语义任务则揭示了精确标注和路径规划方面的明显能力边界。在EASEL-Data上训练的EASEL-9B,相对基础模型提升了6.3%,在所有评估模型中排名第三。
英文摘要
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.