arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29309cond-mat.mtrl-scics.AIphysics.app-ph

评估集成了材料合成工具的基于大语言模型(LLM)的智能体:以原子层沉积为例

Evaluating LLM-based AI agents integrated with materials synthesis tools: the case of atomic layer deposition

Angel Yanguas-Gil

首次发表
浏览论文内容

中文总结 AI 辅助

本研究以原子层沉积为案例,提出了评估集成材料合成工具的基于LLM的AI智能体的实用框架,涵盖多类基准并可推广至其他材料合成技术。

中文摘要 AI 辅助

本研究概述了用于评估基于大语言模型(LLM)的AI模型及材料合成智能体性能的不同策略。在简要介绍当前基于LLM的AI智能体背后的关键技术后,我们总结了材料科学领域、尤其是材料合成场景下评估这些模型的不同方法,重点关注模型直接与实验工具集成的场景。我们讨论的评估策略涵盖知识与推理基准、工具使用基准,以及涉及与实验系统或真实虚拟工具交互的闭环基准。我们以原子层沉积(ALD)为案例研究,强调文献中现有方法如何借鉴材料科学之外的通用方法,以及如何推广到其他材料合成技术。最后,我们提供了一个用于评估材料合成场景下LLM的实用框架。

英文摘要

This work provides an overview of the different strategies that can be used to evaluate the performance of AI models and agents based on large language models (LLMs) for materials synthesis. After providing a brief overview of the key technologies behind the current generation of AI agents based on LLMs, we summarize the different approaches to evaluating these models in the context of materials science and in particular on materials synthesis, with a specific emphasis on scenarios in which the models are directly integrated with experimental tools. We discuss evaluation strategies spanning knowledge and reasoning benchmarks, tool-use benchmarks, and closed loop benchmarks involving the interaction with experimental systems or realistic virtual tools. We use atomic layer deposition (ALD) as a case study, emphasizing how existing approaches in the literature both build from general approaches used beyond materials science and can be generalized to other materials synthesis techniques. Finally, we provide a practical evaluation framework to evaluate LLMs in the context of materials synthesis

发表机构

  • Applied Materials Division, Argonne National Laboratory(阿贡国家实验室应用材料分部)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑