发表机构
Indian Institute of Technology, Delhi(印度理工学院德里分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DISSOLVR提出透明且快速的溶解度预测框架,通过物理描述符实现接近实验极限的精度与OOD泛化,并借助LLM解释流水线提升可解释性。
AI 中文摘要
高保真度的溶解度预测对于药物开发和环境分配至关重要,其中准确的建模必须将分子结构与跨多种化学环境的热力学行为相结合。然而,近期进展主要由深度学习架构主导,这些架构往往以牺牲物理可解释性为代价来换取预测能力。我们通过展示最先进的性能并不需要此类不透明的架构来挑战这一趋势。为此,我们引入了DISSOLVR,一个透明的分子溶解度预测框架。此外,我们进行了全面的文献综述,并针对多种方法开展了基准测试研究。我们证明DISSOLVR接近实验不确定性的偶然极限,并通过结构不变性实现分布外(OOD)泛化,该不变性通过将分子映射到基于物理的描述符而获得。随后,我们提出了一种LLM辅助的事后解释流水线,以弥合符号模型产物与化学基础叙述之间的差距。最后,一项涉及22位专家化学家的比较基准调查显示,专家评估者提供了深刻的见解。
英文摘要
High-fidelity solubility prediction is fundamental to pharmaceutical development and environmental partitioning, where accurate modeling must couple molecular structure with thermodynamic behavior across diverse chemical environments. However, recent advancements have been dominated by deep learning architectures that often sacrifice physical interpretability for predictive power. We challenge this trend by showing that state-of-the-art performance does not require such non-transparent architectures. To address this, we introduce DISSOLVR, a transparent framework for molecular solubility prediction. In addition, we perform a comprehensive literature review and a benchmarking study against various methods. We show that DISSOLVR approaches the aleatoric limit of experimental uncertainty and achieves OOD generalization through structural invariance, derived by mapping molecules to physically-grounded descriptors. Then, we present an LLM-assisted post-hoc explanation pipeline that bridges the gap between symbolic model artifacts and chemically grounded narratives. Finally, a comparative benchmark of a survey involving 22 expert chemists reveals that expert evaluators provide deep insights.
Comments48 pages, 19 tables, 6 figures. Accepted to the 43rd International Conference on Machine Learning (ICML 2026)
Journal refProceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026