AI 中文总结
该研究提出机器学习电荷密度(MLCD)路径,基于含16-96原子的96个小超胞训练集,可准确预测360原子超胞中四个本征缺陷的形成能,其性能优于机器学习原子间势(MLIPs),能大幅降低缺陷预测成本。
AI 中文摘要
第一性原理缺陷计算常受限于抑制镜像相互作用所需的大超胞的成本。机器学习原子间势(MLIPs)提供了另一种替代方案,但训练缺陷MLIPs通常需要数千个结构和数周的数据生成。由于电荷密度是密度泛函理论(DFT)的核心,我们提出一种机器学习电荷密度(MLCD)路径,以更高的数据效率预测缺陷形成能。我们通过整合不同尺寸的小超胞优化训练集,以实现更好的外推,并基于空间电荷密度分析分配它们的比例。仅使用包含16至96个原子的96个超胞作为数据集,MLCD可准确预测360原子超胞中四个本征缺陷的形成能,每个缺陷的平均绝对误差低于0.05 eV。相比之下,在相同数据集上训练的MLIPs的误差可超过1 eV。这些结果表明,电荷密度学习能实现比直接能量-力拟合更稳健的跨尺寸迁移,且混合尺寸数据设计可大幅降低缺陷预测的成本。
英文摘要
First-principles defect calculations are often limited by the cost of the large supercells required to suppress image interactions. Machine-learning interatomic potentials (MLIPs) provide another alternative, but training defect MLIPs typically requires thousands of structures and weeks of data generation. Since charge density is the key to density-functional-theory (DFT), we propose a machine-learning charge density (MLCD) route for predicting defect formation energies with higher data efficiency. We optimize the training set by integrating small supercells of varying sizes for better extrapolation, allocating their proportions based on spatial charge-density analysis. With only 96 supercells containing 16--96 atoms as the dataset, MLCD accurately predicts the formation energies of four intrinsic defects in 360-atom supercells, with defect-wise mean absolute error below 0.05 eV. In contrast, MLIPs trained on the same dataset can err by more than 1 eV. These results show that charge-density learning enables more robust cross-size transfer than direct energy-force fitting and that mixed-size data design can substantially reduce the cost of defect prediction.