AI 中文总结
研究针对机器学习原子间势构建受DFT训练数据集成本限制的问题,通过统计杠杆分数和CUR型采样进行杠杆引导选择,相比随机等基线,能以更小标记比例恢复精度,减少DFT工作量,提升线性ACE模型数据效率。
AI 中文摘要
机器学习原子间势(MLIPs)的构建常受生成大型密度泛函理论(DFT)训练数据集成本的限制。对于像ASSYST这样系统生成的结构池,关键问题是需标记多少构型才能实现可靠精度。本文评估用于训练线性原子簇展开(ACE)势的基于几何的无标记子集选择。利用统计杠杆分数和CUR型采样,在受控迭代协议下,将杠杆引导选择与随机、基于能量和基于力的基线进行比较。以元素Al为主要基准,用Cu和Al-Cu合金进行转移验证。杠杆引导子集使用比随机采样小得多的标记比例(约3%-40%)就能恢复平台级能量和力精度,相当于所研究系统的DFT标记有效减少2-3倍。在合金测试中,一旦包含足够化学多样性,各策略的缺陷能量学相当,且杠杆选择在减少训练规模时保持有竞争力的精度。这些结果表明,描述符空间引导的无标记子采样可显著减少在ASSYST结构池上训练的线性ACE模型的DFT工作量,且不降低缺陷级保真度。
英文摘要
The construction of machine-learned interatomic potentials (MLIPs) is often limited by the cost of generating large density-functional-theory (DFT) training datasets. For systematically generated structure pools such as ASSYST, a central practical question is how many configurations must be labeled to achieve reliable accuracy. Here we assess geometry-based, label-free subset selection for training linear Atomic Cluster Expansion (ACE) potentials. Using statistical leverage scores and CUR-type sampling, we compare leverage-guided selection against random, energy-based, and force-based baselines under controlled iterative protocols. Elemental Al provides the primary benchmark, with Cu and Al-Cu alloys used for transfer validation. Leverage-guided subsets recover plateau-level energy and force accuracy using substantially smaller labeled fractions (approximately 30-40%) than random sampling, corresponding to an effective 2-3x reduction in DFT labeling for the systems studied. In alloy tests, defect energetics remain comparable across strategies once sufficient chemical diversity is included, while leverage selection maintains competitive accuracy at reduced training size. These results demonstrate that descriptor-space-guided, label-free subsampling can significantly reduce DFT workload for linear ACE models trained on ASSYST structure pools without degrading defect-level fidelity.
Comments16 pages, 6 figures