发表机构
Science and Technology Facilities Council(科学与技术设施委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较了以基态电子密度与分子几何构型作为机器学习模型输入预测吸收光谱的效果,发现基于密度的卷积神经网络显著优于几何图模型,验证相关性达0.9926,残差减少约64%。
AI 中文摘要
分子光学吸收光谱提供了电子结构的直接探针,广泛用于分子识别、光物理行为解释以及光谱实验规划。然而,使用第一性原理激发态方法计算吸收光谱在计算上要求很高,至少与基态计算相比,这限制了其在大型分子集合中的常规应用。机器学习(ML)代理可以降低这一成本并实现快速光谱预测。然而,其性能在很大程度上取决于分子信息的表示方式。在此,我们比较了使用基态电子密度与分子几何构型作为机器学习模型输入来预测吸收光谱的效果,训练集包含从QM7数据集中选取的6874个分子。对于每个分子,使用密度泛函理论(DFT)计算密度,并使用线性响应(LR)含时密度泛函理论(TDDFT)计算吸收光谱。利用基态密度作为机器学习模型的输入,其动机源于Hohenberg-Kohn和Runge-Gross定理,以及基态密度编码了关于成键、电荷局域化和电子离域化的信息。因此,与几何构型相比,它可能是机器学习模型更明智的起点,因为它有效地解耦了基态的化学性质。我们测试的问题是,使用密度的好处是否超过需要额外进行单点DFT计算密度的(非禁止性)代价。我们发现,基于密度的卷积神经网络实现了0.9926的验证相关性,而最佳基于几何构型的图模型为0.9795,残差去相关减少了约64%。
英文摘要
Molecular optical absorption spectroscopy provides a direct probe of electronic structure and is widely used for molecular identification, interpretation of photophysical behaviour, and planning of spectroscopy experiments. Calculating the absorption spectra using first-principle excited-state methods, however, is computationally demanding, at least compared to ground-state calculations, which limits their routine application across large molecular sets. Machine-learning (ML) surrogates can reduce this cost and allow rapid spectral prediction. However, their performance depends strongly on how molecular information is represented. Here, we compare using the ground-state electron density versus the molecular geometry as inputs to a ML model for predicting absorption spectra, for a training set of 6874 molecules selected from the QM7 dataset. For each of these molecules, the density was calculated using density functional theory (DFT) and the absorption spectrum was calculated using linear-response (LR) time-dependent DFT (TDDFT). Utilizing the ground-state density as the input to the ML model is motivated by the Hohenberg-Kohn and Runge-Gross theorems, and the fact that the ground-state density encodes information about bonding, charge localisation, and electronic delocalisation. Hence, it may be a more judicious starting point for the ML model compared to the geometry, as it effectively decouples the chemistry of the ground-state. The question we test is whether the benefits of using the density outweigh the (notprohibitive) penalty of requiring an additional single-point DFT calculation for the density. We find that the density-based convolutional neural network achieves a validation correlation of 0.9926, compared with 0.9795 for the best geometry-based graph model, reducing the residual decorrelation, by approximately 64%.