arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05877cs.LG

预算依赖的基于覆盖与基于响应的训练集选择在机器学习原子间势中的交叉

A budget-dependent crossover between coverage- and response-based training-set selection for machine-learned interatomic potentials

  • Digital Discovery(数字发现)

机构由 AI 辅助整理,请以论文原文为准。

Jia Bi, Alin-Marin Elena

AI总结:

本文通过预算分辨比较,发现训练集选择中基于响应的见证方法在20%预算时优于结构覆盖,而在低预算时覆盖更优,确立了数据预算为原子势训练集选择的关键变量。

AI中文摘要:

为机器学习原子间势选择紧凑训练集需要决定是保留结构多样性还是针对模型存在分歧的配置。更好的选择可能取决于保留的数据量,这使得在单一训练集规模下的比较不够充分。在此,我们通过预算分辨比较在GAP-20 Carbon和合并修订的MD17上重训练的MACE模型,将选择标准与预测精度联系起来。我们将结构覆盖与一个响应引导的选择器进行比较,该选择器针对覆盖训练模型与全数据参考之间的分歧。这个回顾性响应见证测试了模型分歧在压缩已标记池中的价值。在5%时,覆盖在两个数据集的四个力端点上的绝对偏差均小于随机采样相对于全数据误差的偏差。见证在1%和5%时的偏差大于覆盖,但在20%时顺序反转。在20%时,见证选择的模型相对于覆盖,直接留出力误差降低了0.46--5.89%,所有八个配对的训练种子区间都倾向于见证。六个误差低于全数据参考。平均力误差降低为0.164--0.167 meV Å⁻¹,尾部端点和掩码端点有更大增益。补充分析表明,学习到的相似性保持了覆盖排序,而按冻结模型误差选择则给出比嵌入覆盖更高的误差。这些发现确立了保留数据预算作为原子训练集选择中的决定性变量,并提供了响应引导压缩何时优于结构覆盖的直接测试。

英文摘要:

Selecting compact training sets for machine-learned interatomic potentials requires deciding whether to preserve structural diversity or target configurations on which models disagree. The better choice can depend on how much data is retained, making a comparison at one training-set size insufficient. Here we link selection criteria to prediction accuracy through a budget-resolved comparison of retrained MACE models on GAP-20 Carbon and pooled revised MD17. Structural coverage is compared with a response-guided selector that targets disagreement between a coverage-trained model and a full-data reference. This retrospective response witness tests the value of model disagreement for compressing an already labelled pool. At 5\%, coverage gives smaller absolute deviations from the full-data error than random sampling across four force endpoints in both datasets. The witness has larger deviations than coverage at 1\% and 5\%, but the ordering reverses at 20\%. At 20\%, witness-selected models also lower direct held-out force errors by 0.46--5.89\% relative to coverage, with all eight paired training-seed intervals favouring the witness. Six errors fall below the full-data reference. Mean force-error reductions are 0.164--0.167~meV~$\textÅ^{-1}$, with larger gains for tail and masked endpoints. Complementary analyses show that learned similarity preserves the coverage ranking, while selecting by frozen-model error gives higher error than embedding coverage. These findings establish retained-data budget as a deciding variable in atomistic training-set selection and provide a direct test of when response-guided compression improves on structural coverage.

↑