发表机构
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院自动化研究所; 中国科学院大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
S3C-LLM是一种技能-代码引导型智能体LLM,通过检索光谱学技能、执行分析代码整合证据与约束来解析分子结构,性能优于同类模型且训练语料更少。
AI 中文摘要
谱图结构解析是分子分析的核心,但近期基于大语言模型(LLM)的方法大多将其表述为直接从谱图生成SMILES的任务。尽管该范式可利用配对的谱图数据,但未明确建模光谱学家使用的分析流程,如诊断峰解读、片段推理、分子式约束及化学一致性检查。本文提出S3C-LLM,一种用于谱图到结构解析的技能引导且代码落地的智能体LLM。S3C-LLM不直接预测分子,而是检索模态特定的光谱学技能,执行分析代码以在输入谱图上实例化这些技能,整合所得的峰级证据与分子式约束后再生成SMILES。具体而言,本文贡献了自演化光谱学技能库、思维增强的技能-代码轨迹构建流程,以及两阶段训练策略,通过监督微调(SFT)及本文提出的步级强化学习(RL)来训练Qwen3-4B。在各类基准测试上的实验表明,S3C-LLM在各类谱图上的表现始终优于当前通用LLM及谱图专用模型,且所用训练语料不足SpectraLLM的1/10。
英文摘要
Spectroscopic structure elucidation is central to molecular analysis, but recent Large Language Model (LLM)-based methods mostly formulate it as direct spectrum-to-SMILES generation. Although this paradigm can leverage paired spectral data, it does not explicitly model the analytical workflow used by spectroscopists, such as diagnostic peak interpretation, fragment reasoning, formula constraints, and chemical consistency checking. In this paper, we introduce S3C-LLM, a skill-guided and code-grounded agentic LLM for spectrum-to-structure elucidation. Rather than directly predicting a molecule, S3C-LLM retrieves modality-specific spectroscopy skills, executes analysis code to instantiate these skills on the input spectra, and integrates the resulting peak-level evidence and formula constraints before generating SMILES. Specifically, we contribute a self-evolving spectroscopy skill library, a thinking-augmented skill-code trajectory construction pipeline, and a two-stage training strategy that teaches Qwen3-4B through supervised fine-tuning (SFT) followed by our proposed step-level reinforcement learning (RL). Experiments on diverse benchmarks show that S3C-LLM consistently outperforms current general LLMs and spectrum-specific models across spectra, while using less than 1/10th of SpectraLLM's training corpus.
CommentsAccepted by EMNLP 2026 Findings