arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24513cond-mat.mtrl-sci

基于化学语言模型的电子态密度的结构无关预测

Structure-Agnostic Prediction of the Electronic Density of States with a Chemical Language Model

Ivan D. Rubtsov, Ivan V. Dudakov, Vadim V. Korolev

AI总结:

该研究提出化学语言模型DOSSIER,无需晶体结构即可直接由元素组成预测电子态密度,在基准测试及合金筛选中表现出良好性能,为未合成化合物的DOS预测提供了结构无关的新方法。

AI中文摘要:

电子态密度(DOS)通常通过弛豫晶体结构计算,但该结构无法获取未合成或未编目的化合物。本文引入DOSSIER(全称:基于化学计量比与编码器表征的态密度模型),一种可直接将元素组成映射至该谱的化学语言模型。其编码器通过从通用机器学习原子间势进行跨模态知识蒸馏完成预训练;当仅用1000个训练样本时,该迁移过程使误差降低11%。在Mat2Spec基准测试中,DOSSIER的平均绝对误差为3.76态·eV⁻¹,而最优结构感知模型为3.64;在扩展的Materials Project数据集上,预测谱可准确得到带隙和d带描述符。针对11977种二元及13251种五元高熵合金组成,筛选出与NiPt₃的d投影DOS相似的体系,结果显示已知氧还原电催化剂位于排名前列。

英文摘要:

The electronic density of states (DOS) is conventionally computed from a relaxed crystal structure, which is unavailable for compounds that have been neither synthesized nor cataloged. Here we introduce DOSSIER ($\textbf{D}$ensity $\textbf{o}$f $\textbf{S}$tates from $\textbf{S}$to$\textbf{i}$chiometry with $\textbf{E}$ncoder $\textbf{R}$epresentations), a chemical language model that maps elemental composition directly to this spectrum. The encoder is pretrained by cross-modal knowledge distillation from a universal machine-learning interatomic potential; the transfer lowers the error by 11% when only 1,000 training examples are available. On the Mat2Spec benchmark, DOSSIER reaches a mean absolute error of 3.76 states eV$^{-1}$ against 3.64 for the best structure-aware model; on an extended Materials Project dataset, the predicted spectra yield band gaps and $\textit{d}$-band descriptors with useful accuracy. Screening 11,977 binary and 13,251 five-component high-entropy alloy compositions for a $\textit{d}$-projected DOS resembling that of NiPt$_{3}$ places known oxygen reduction electrocatalysts near the top of the ranking.

补充信息

↑