arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在训练和微调机器学习力场时使用更少标签实现全数据精度

Full-data accuracy with fewer labels for training and fine-tuning machine-learning force fields

Sheng Bi, Yi-Ze Wang, Jun Cheng

arXiv 2607.14486首次发表:更新:

发表机构

College of Materials; Xiamen University; State Key Laboratory of Physical Chemistry of Solid Surfaces; iChEM, College of Chemistry and Chemical Engineering; Laboratory of AI for Electrochemistry (AI4EC); IKKEM(材料学院; 厦门大学; 固体表面物理化学国家重点实验室; iChEM,化学与化工学院; 电化学人工智能实验室; IKKEM)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对机器学习力场构建训练集的瓶颈问题,提出基于最后一层投影回归的主动学习工作流程,能识别高价值训练集,在基础模型微调和迭代电解质微调中表现出色,实现用更少标签达到全数据精度,提供可扩展的不确定性量化策略。

AI 中文摘要

机器学习力场(MLFFs)仅在其训练分布附近可靠,这使得构建多样的训练集成为从零训练和基础微调工作流程的主要瓶颈。主动学习可降低成本,但标准模型委员会不确定性对基础MLFFs不实用。本文提出基于最后一层投影回归(LLPR)的主动学习工作流程。在多个系统中,LLPR能识别紧凑、高价值训练集,仅用少量电子结构标签就能恢复全数据精度。在基础模型微调中,LLPR选择的配置用比随机选择少得多的标签达到全池微调上限。在迭代电解质微调中,LLPR能在DFT标记前检测非物理局部配位,提供绝对力误差阈值并实现学习循环自动终止。

英文摘要

Machine-learning force fields (MLFFs) are reliable only near their training distribution, making efficient construction of diverse training sets a major bottleneck for both train-from-scratch and foundation fine-tuning workflows. Active learning can reduce this cost, but standard model-committee uncertainty is impractical for foundation MLFFs because each committee member requires a separate fine-tuning run. We present an active-learning workflow based on last-layer-projection regression (LLPR), a forward-pass-cheap per-configuration uncertainty estimator. Across molecular, condensed-phase, and electrolyte systems, LLPR identifies compact, high-value training sets that recover full-data accuracy using only a small fraction of electronic-structure labels. In foundation-model fine-tuning, LLPR-selected configurations reach the full-pool fine-tuning ceiling with substantially fewer labels than random selection. In iterative electrolyte fine-tuning, LLPR detects unphysical local coordination before DFT labelling, provides an absolute force-error threshold, and enables automatic termination of the learning loop. The resulting models reproduce reference density and ion-coordination structure, providing a scalable uncertainty-quantification strategy across MLFF training regimes.

Comments23 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑