基于多模态光谱数据的有机结构假设与细化学习
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
- State Key Laboratory of Advanced Waterproof Materials, School of Materials Science and Engineering, Peking University(先进防水材料国家重点实验室,材料科学与工程学院,北京大学)
- School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(电子与计算机工程学院,北京大学深圳研究生院)
- Peng Cheng Laboratory(鹏城实验室)
- IBS Center for Algorithmic and Robotized Synthesis (CARS), UNIST(算法与机器人化合成中心(CARS),UNIST)
- Department of Chemistry, UNIST(化学系,UNIST)
- School of Advanced Materials, Peking University Shenzhen Graduate School(先进材料学院,北京大学深圳研究生院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对从光谱数据确定分子结构的难题,提出假设细化范式,构建QM9SPIN数据集,引入SpectroMol和MS - Mol2Mol,集成系统在模拟基准上准确率达93.8%,能适应实验光谱并改进预测,为有机结构解析提供可扩展途径。
AI中文摘要:
从光谱数据确定分子结构具有根本挑战性,因为反问题本质上是欠定的。本文提出将自动结构解析作为一种可扩展的假设细化范式,紧密整合光谱证据与大规模分子先验。构建了QM9SPIN数据集,引入了SpectroMol模型和MS - Mol2Mol分子生成器。集成系统在模拟基准上实现了93.8%的top - 1准确率,能有效从模拟光谱适应到实验光谱,并通过质量引导细化改进实验预测,为自动化、数据驱动的有机结构解析建立了可扩展途径。
英文摘要:
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.