arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.19816physics.chem-phcs.LGphysics.comp-phphysics.data-an

基于多模态光谱数据的有机结构假设与细化学习

Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

  • State Key Laboratory of Advanced Waterproof Materials, School of Materials Science and Engineering, Peking University(先进防水材料国家重点实验室,材料科学与工程学院,北京大学)
  • School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(电子与计算机工程学院,北京大学深圳研究生院)
  • Peng Cheng Laboratory(鹏城实验室)
  • IBS Center for Algorithmic and Robotized Synthesis (CARS), UNIST(算法与机器人化合成中心(CARS),UNIST)
  • Department of Chemistry, UNIST(化学系,UNIST)
  • School of Advanced Materials, Peking University Shenzhen Graduate School(先进材料学院,北京大学深圳研究生院)

机构由 AI 辅助整理,请以论文原文为准。

Chengchun Liu, Zhiyuan Yan, Li Yuan, Hao Li, Boxuan Zhao, Yonghong Tian, Bartosz A. Grzybowski, Fanyang Mo

AI总结:

该研究针对从光谱数据确定分子结构的难题,提出假设细化范式,构建QM9SPIN数据集,引入SpectroMol和MS - Mol2Mol,集成系统在模拟基准上准确率达93.8%,能适应实验光谱并改进预测,为有机结构解析提供可扩展途径。

AI中文摘要:

从光谱数据确定分子结构具有根本挑战性,因为反问题本质上是欠定的。本文提出将自动结构解析作为一种可扩展的假设细化范式,紧密整合光谱证据与大规模分子先验。构建了QM9SPIN数据集,引入了SpectroMol模型和MS - Mol2Mol分子生成器。集成系统在模拟基准上实现了93.8%的top - 1准确率,能有效从模拟光谱适应到实验光谱,并通过质量引导细化改进实验预测,为自动化、数据驱动的有机结构解析建立了可扩展途径。

英文摘要:

Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.

补充信息

↑