arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CompMat-Bench:计算材料科学AI智能体基准测试

CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson

arXiv 2610.00636首次发表:更新:

发表机构

Rice University(莱斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CompMat-Bench提出94个计算材料科学任务的基准,通过预复现模拟步骤评估AI智能体,避免昂贵计算,支持多种评估条件,揭示智能体性能差异及主要失败源于科学错误。

AI 中文摘要

在科学研究任务上评估AI智能体受到基础实验或计算所需时间和资源的限制。在计算材料研究中,跨智能体和试验重复相同的昂贵模拟可能使评估变得不切实际。我们引入了CompMat-Bench,一个包含94个任务的基准测试,这些任务源自最近发表的计算材料研究,每个任务要求智能体完成朝向实现研究科学目标的一步。我们预先复现研究步骤,并评估智能体在准备昂贵模拟的输入和分析输出方面的能力,从而在评估过程中避免昂贵的模拟。复现的输入和结果作为使用固定规则(无需LLM评判)对智能体进行评分的真实依据。该基准支持四种评估条件:单一任务和由相关任务组成的工作流,每种条件均提供完整或简化的方法指导。在完整指导下的单一任务中,基于三个LLM的智能体展示了完成单个材料研究步骤的能力,在94个任务中的通过率为66.0%至90.4%。较长的工作流和简化的指导都可能限制智能体的性能,但对不同智能体的影响方式不同:它们降低了较弱智能体的通过率,而最强的智能体仅在长工作流与简化指导结合时才失败。失败分析将大多数失败归因于科学错误而非软件使用错误。CompMat-Bench为在真实材料研究步骤上比较智能体和分析智能体失败模式提供了基础。

英文摘要

Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study's scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents with fixed rules, without an LLM judge. The benchmark supports four evaluation conditions: single tasks and workflows composed of related tasks, each with full or reduced methodological guidance. With full guidance on single tasks, agents based on three LLMs demonstrate the ability to complete individual materials research steps, with pass rates of 66.0-90.4% across 94 tasks. Both longer workflows and reduced guidance can limit agent performance, but in different ways for different agents: they lower the pass rates of the weaker agents, whereas the strongest agent falls only when a long workflow is combined with reduced guidance. Failure analysis attributes most failures to scientific errors rather than to errors in software usage. CompMat-Bench provides a basis for comparing agents on the steps of real materials research and for analyzing agent failure modes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑