arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多粒度理由引导的分子大语言模型用于属性预测

Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee

arXiv 2608.10480首次发表:更新:

AI 中文总结

本研究提出多粒度理由引导的分子大语言模型MR-MoL,将GNN导出的子结构属性归因作为证据,在8项MoleculeNet任务上取得通用模型最佳结果,缩小了与专用模型的差距。

AI 中文摘要

大语言模型(LLM)被广泛应用于化学任务,如分子属性预测,这是药物发现的基础。分子LLM通过多种模态表示分子,尤其是1维SMILES序列或2维分子图,两者均隐含编码分子信息,因此单个子结构的贡献仍不明确。检索与增强方法会从外部来源添加上下文,但化学家推理的线索是驱动属性升降的内部子结构。我们提出MR-MoL,一种多粒度理由引导的分子LLM,直接提供此类证据:微调后的GNN通过掩码对每个子结构打分,最具影响力的子结构被序列化为带排名和方向标签的理由,LLM将其与SMILES序列、分子图一同读取。该理由涵盖三种粒度:带有侧链的Murcko骨架、BRICS片段和官能团。据我们所知,这是首个将GNN导出的属性归因作为证据提供给LLM用于属性预测的方法。在8项MoleculeNet任务上,MR-MoL在通用模型中取得整体最佳结果,且缩小了与针对各任务调优的专用模型的差距。5项诊断进一步证实模型确实读取了理由,而非仅得益于其存在;理由的方向、排名和子结构均会影响预测,且其归因可复现已知的结构-属性关系。

英文摘要

Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular graph. Both encode molecular information implicitly, so the contribution of individual substructures remains opaque. Retrieval and augmentation methods add context, but from external sources. However, the cues chemists reason over are the internal substructures that drive a property up or down. We propose MR-MoL, a multi-granular rationale-guided molecular LLM that supplies this evidence directly. A fine-tuned GNN scores each substructure through masking, and the most influential ones are serialized as a ranked, direction-tagged rationale that the LLM reads alongside the SMILES sequence and molecular graph. The rationale spans three levels of granularity: Murcko scaffolds with their side chains, BRICS fragments, and functional groups. This is, to our knowledge, the first method to expose GNN-derived attributions to an LLM as evidence for property prediction. On eight MoleculeNet tasks, MR-MoL achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task. Five diagnostics further confirm that the model reads the rationale rather than merely benefiting from its presence. Its direction, rank, and substructure each shape the prediction, and its attributions reproduce known structure-property relationships.

Comments16 pages, 5 figures, 18 tables. Code: https://github.com/skku-aihclab/MR-MoL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑