AI 中文总结
本研究提出HADES混合机器学习模型,结合深度集神经网络与Stokes-Einstein方程,预测任意组分和温度下液体混合物的自扩散系数,仅需SMILES结构和纯组分粘度,在2526个数据点上显著优于基准方法。
AI 中文摘要
自扩散系数是分子迁移率的关键描述符,然而实验数据仍然稀缺,这凸显了对可靠预测方法的需求。在先前的工作中,我们引入了混合增强Stokes-Einstein(ESE)模型,该模型通过将Stokes-Einstein方程与机器学习(ML)相结合,推进了纯溶剂中无限稀释溶质自扩散系数物理一致预测的最新技术水平。在此,我们将该方法扩展到浓度依赖的自扩散系数以及多组分溶剂,提出了HADES。这种混合架构利用深度集神经网络,在单一框架内连接纯组分预测和混合物预测。HADES能够预测任意组分数量、任意组成和温度下液体混合物中的自扩散系数。唯一需要的输入是组分的SMILES编码分子结构和纯组分粘度,这使得该方法具有广泛的适用性。在一个包含600个体系、2526个数据点的综合数据集上进行训练和评估,HADES显著优于基准预测方法。训练好的模型及其源代码完全公开,并且该应用可通过交互式网站(此https URL)获取。
英文摘要
Self-diffusion coefficients are key descriptors of molecular mobility, yet experimental data remain scarce, highlighting the need for reliable prediction methods. In previous work, we introduced the hybrid Enhanced Stokes-Einstein (ESE) model, which advanced the state of the art in the physically consistent prediction of self-diffusion coefficients of solutes at infinite dilution in pure solvents by integrating the Stokes-Einstein equation with machine learning (ML). Here, we extend this approach to concentration-dependent self-diffusion coefficients and multicomponent solvents with HADES. This hybrid architecture leverages a deep-set neural network to connect pure-component and mixture prediction within a single framework. HADES predicts self-diffusion coefficients in liquid mixtures with any number of components at any composition and temperature. The only required inputs are SMILES-encoded molecular structures of the components and the pure-component viscosities, making the method broadly applicable. Trained and evaluated on a comprehensive dataset of 2526 data points for 600 systems, HADES significantly outperforms benchmark prediction methods. The trained model and its source code are fully disclosed, and the application is available via an interactive website https://ml-prop.mv.rptu.de/.