使用基于图的工具改进小语言模型中的分子性质预测
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
浏览论文内容
中文总结 AI 辅助
研究针对小语言模型在分子性质预测时的结构盲目性问题,提出模块化上下文增强提示框架,通过GNN专家模型和提取子图来辅助预测,在MUTAG和Tox21数据集上实验表明该方法能显著提升预测准确性,验证了基序相关性,但与专门GNN模型仍有差距。
中文摘要 AI 辅助
小语言模型在从SMILES字符串进行零样本分子性质预测方面展现出潜力,但常因结构盲目性而受限,因为序列表示未充分体现关键图拓扑线索。我们提出模块化上下文增强提示框架,在推理时启用智能工具使用:训练好的GNN专家模型提供有信心的预测提示,GNN提取特定实例的解释性子图。我们在MUTAG和Tox21上评估了三种常用小语言模型,在五种提示配置下从仅SMILES到使用所有可用工具。在两个数据集上,用图衍生上下文丰富提示可显著提高准确性,Tox21上相对改进常超25%,最高达74%。我们还通过基于必要性的边删除干预验证了提取基序的功能相关性。尽管有改进,但与专门的GNN模型仍有差距,凸显了文本条件推理对分子结构的价值和局限性。
英文摘要
Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural blindness because sequence representations under-specify key graph-topological cues. We propose a modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hint with confidence, and a GNN extracts an instance-specific explanatory subgraph (e.g., a subgraph SMILES and an accompanying explanatory paragraph). We evaluate three commonly used SLMs on MUTAG and Tox21 under five prompting configurations ranging from SMILES-only to using all available tools at hand. Across two datasets, enriching prompts with graph-derived context yields substantial accuracy gains, often exceeding 25% relative improvement and up to 74% on Tox21. We further validate the functional relevance of the extracted motifs via a necessity-based edge-drop intervention. Despite the observed gains, a persistent gap remains to specialized GNN models, highlighting both the value and limits of text-conditioned reasoning for molecular structure.
发表机构
- Institute of Informatics and Telecommunications, National Center for Scientific Research “Demokritos(信息与电信研究所,国家科学研究中心“德谟克利特”)
机构由 AI 辅助整理,请以论文原文为准。