AI 中文总结
提出DiffGCMS,结合扩散模型与LLM重排序,从GC-EI-MS谱推断分子结构,提升准确性与可解释性。
AI 中文摘要
GC-EI-MS是分析复杂样品中挥发性与半挥发性化合物的重要技术。然而,传统方法严重依赖参考谱库匹配,限制了其识别谱库中缺失化合物以及直接从碎片信息推断完整分子结构的能力。在此,我们提出DiffGCMS,一种谱条件离散图扩散模型,用于从GC-EI-MS进行从头结构解析,并进一步开发了一个将DiffGCMS与大型语言模型(LLM)的第二阶段推理相结合的框架。在第一阶段,DiffGCMS从输入谱生成候选分子结构;在第二阶段,LLM利用质谱信息验证、修复和重新排序候选结构,并提供碎片离子峰的合理解释。该框架能够为参考谱库中缺失的化合物生成合理的分子结构,并提供可追溯的证据支持其决策。在包含NIST 20中13,696个谱的测试集上,生成模型实现了Acc@1和Acc@10分别为6.01%和15.76%。在包含不超过10个重原子的分子的测试子集上,LLM辅助的分子图修复和重新排序将Acc@1从21.28%提高到21.95%,Acc@10从46.91%提高到47.99%,候选有效性从91.04%提高到100%。这些结果表明,谱感知的后处理可以纠正生成模型产生的错误,同时为最终排序提供可审计和可追溯的解释。
英文摘要
GC--EI--MS is an important technique for analyzing volatile and semivolatile compounds in complex samples. However, conventional methods rely heavily on reference spectral library matching, limiting their ability to identify compounds absent from these libraries and to infer complete molecular structures directly from fragmentation information. Here, we present DiffGCMS, a spectrum-conditioned discrete graph diffusion model for de novo structure elucidation from GC--EI--MS, and further develop a framework that integrates DiffGCMS with second-stage reasoning by a large language model (LLM). In the first stage, DiffGCMS generates candidate molecular structures from input spectra; in the second stage, the LLM uses mass spectral information to validate, repair, and rerank the candidates and provides interpretable analysis of fragment-ion peaks. This framework can generate plausible molecular structures for compounds absent from reference spectral libraries and provide traceable evidence supporting its decisions. On a test set comprising 13,696 spectra from NIST 20, the generative model achieved Acc@1 and Acc@10 of 6.01\% and 15.76\%, respectively. On the test subset containing molecules with no more than 10 heavy atoms, LLM-assisted molecular graph repair and reranking increased Acc@1 from 21.28\% to 21.95\%, Acc@10 from 46.91\% to 47.99\%, and candidate validity from 91.04\% to 100\%. These results demonstrate that spectrum-aware postprocessing can correct errors produced by the generative model while providing auditable and traceable explanations for the final ranking.
Comments4 figures