arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MolSC:利用取代基贡献增强大语言模型中的细粒度分子理解

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee

arXiv 2609.23073首次发表:更新:

发表机构

Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出取代基贡献数据集MolSC及基准MolSC-Bench,通过训练提升分子大语言模型对细粒度结构-性质关系的预测能力,实验证明该方法显著优于现有模型。

AI 中文摘要

自然语言处理的最新进展催生了分子大语言模型(LLMs),这些模型在多种化学任务中表现出强大的性能。然而,它们仍然难以捕捉细粒度的结构-性质关系,特别是局部的微小修饰如何改变分子的行为。为了解决这一局限性,我们引入了MolSC,一个取代基贡献数据集,定义为将特定取代基连接到分子骨架时所引起的性质变化。该数据集从人工标注的生物活性记录中整理而成,涵盖结构警示毒性、靶点特异性生物活性和物理化学描述符,并包含181K个取代基级别的训练示例。我们进一步提出了MolSC-Bench,一个留出评估基准,包含1,541个示例,这些示例在骨架、取代基和分子层面上与MolSC不重叠。我们的实验表明,现有的分子LLMs以及如GPT-5.2和Gemini-3-Flash等强大的专有模型在取代基贡献预测方面表现出有限的可靠性。相比之下,在MolSC上训练显著提高了这一能力,并在多种下游分子任务中取得了强劲的性能。这些结果突显了取代基贡献学习是细粒度分子理解的关键组成部分。

英文摘要

Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.

CommentsAccepted to EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑