arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19790cs.AI

大语言模型作为有限池材料优化的获取策略:一项对照研究

LLMs as Acquisition Policies for Finite-Pool Materials Optimization: A Controlled Study

Dino-Rober Demir, Florian Le Bronnec, Rio Yokota

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探究开放权重大语言模型能否作为有限池材料优化的独立获取策略,经对照实验发现其无需任务特定训练即可提供有用获取信号,虽可靠性受多因素影响但显示出应用潜力。

中文摘要 AI 辅助

发现具有理想特性的材料通常需要在庞大的候选空间中搜索,而实验或计算评估的成本仍然很高。主动学习通过利用先前的观测结果选择下一个要评估的候选对象来应对这一挑战,通常是通过概率代理模型实现的。我们研究开放权重大语言模型(LLMs)是否可以在该场景中作为独立的获取策略。我们在四种回顾性有限池材料优化任务中,针对不同的候选呈现策略评估了五个LLMs,并将它们与随机选择和传统高斯过程方法进行比较。LLM策略通常比随机选择在更少的迭代中达到全局最优,表明它们无需针对特定任务进行训练即可提供有用的获取信号。它们相对于高斯过程方法的表现参差不齐:传统获取在大多数任务上表现更好,而LLMs在某些设置中与它相当或优于它。性能在任务、模型、初始化和候选呈现之间差异显著,没有任何LLM方法在所有任务上表现最佳。总体而言,开放权重LLMs作为有限池材料搜索的获取策略显示出潜力,尽管其可靠性仍然对任务以及候选和科学背景的呈现方式敏感。

英文摘要

Discovering materials with desirable properties often requires searching large candidate spaces while experimental or computational evaluations remain costly. Active learning addresses this challenge by using previous observations to select which candidate to evaluate next, typically through probabilistic surrogate models. We investigate whether open-weight large language models (LLMs) can serve as standalone acquisition policies in this setting. We evaluate five LLMs across four retrospective finite-pool materials optimization tasks under different candidate-presentation strategies and compare them with random selection and conventional Gaussian-process methods. LLM policies generally reach the global optimum in fewer iterations than random selection, indicating that they provide a useful acquisition signal without task-specific training. Their performance relative to Gaussian-process methods is mixed: conventional acquisition performs better on most tasks, while LLMs match or outperform it in some settings. Performance varies substantially across tasks, models, initializations, and candidate presentations, with no LLM approach performing best across all tasks. Overall, open-weight LLMs show potential as acquisition policies for finite-pool materials search, although their reliability remains sensitive to the task and to how candidates and scientific context are presented.

发表机构

  • RIKEN(理化学研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑