arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MatToolBench:在真实材料科学工作流中基准测试多模态智能体

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, Lu Chen

arXiv 2609.37053首次发表:更新:

发表机构

Shanghai Jiao Tong University; Suzhou Laboratory; Shanghai Artificial Intelligence Laboratory; State Key Laboratory for General Artificial Intelligence, BIGAI; Shanghai Innovation Institution; Jiangsu Key Lab of Language Computing(上海交通大学; 苏州实验室; 上海人工智能实验室; 北京通用人工智能研究院通用人工智能全国重点实验室; 上海创新研究院; 江苏省语言计算重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MatToolBench是首个针对专业材料科学软件的多模态GUI智能体真实环境基准,包含204个任务,揭示通用基准性能无法迁移至科学工作流,最佳模型GUI任务成功率仅25%。

AI 中文摘要

多模态GUI智能体在通用软件基准测试上取得了令人瞩目的成果,但其操作专业科学软件的能力在很大程度上仍未得到探索。在材料科学领域,稀疏的领域特定网络数据、专业化的界面以及隐性的工作流惯例造成了通用预训练难以弥合的盲区。我们提出了MatToolBench,这是首个用于评估多模态GUI智能体在专业材料科学软件上表现的真实环境基准,包含在Windows 11虚拟机中执行的10个工具上的204个任务,涵盖三种模态:GUI操作、OriginPro脚本编写和基于代码的数据库查询。每个任务由领域专家分解为细粒度的子标准,从而实现可解释的部分评分;我们多层级评估流程中的GUI组件平均F1达到0.98。对于OriginPro图形生成任务,我们进一步进行了人类与LLM一致性研究,以验证使用多模态评判器进行次级美学评估的有效性。我们的实验表明,在通用基准上的强劲表现并不能迁移到专业科学工作流中,且这一差距不仅仅是视觉定位问题:失败源于领域特定的操作知识、科学软件预训练覆盖的稀疏性、跨工具工件交接的薄弱环节以及仅通过视觉暴露的关键状态。即使是最佳模型在GUI任务上也仅达到25%的成功率,在代码任务上达到45%。因此,MatToolBench作为一个具有挑战性的诊断基准和真实环境测试平台,适用于数据稀缺、知识密集的科学工作流。

英文摘要

Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.

Comments25 pages, 15 figures. Mei Wu and Rui Xie contributed equally. Bo Chen and Lu Chen are corresponding authors. Project page: https://mattoolbench.github.io/ ; code: https://github.com/meiwu5/MatToolBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑