arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于SMILES与自然语言联合表示学习的双语义化学嵌入器

Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language

David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Ranković, Philippe Schwaller

arXiv 2608.03855首次发表:更新:

AI 中文总结

该研究提出双语义化学嵌入模型CheMatE,通过两阶段训练将SMILES与自然语言联合表示,在分子性质预测等下游任务上表现具竞争力。

AI 中文摘要

Transformer模型彻底改变了自然语言处理(NLP)领域,而基于文本的分子表示(如SMILES)已成功将这些架构扩展到化学领域。然而,领域自适应预训练常导致模型过拟合于化学语法,灾难性地遗忘其基础语义能力。为应对这一挑战,我们提出CheMatE——一种面向化学的嵌入模型,它在同一表示空间中联合捕捉分子结构与领域特定自然语言。该模型基于ModernBERT主干构建,通过两阶段训练流程学习双语义表示:首先是延续性掩码语言建模(MLM),随后是通过多负排序损失(MNRL)实现的Matryoshka对比学习阶段。首先,我们在一个新的大规模语料库上对模型进行MLM训练,该语料库由从FineWeb和ChemPile构建并整理的带SMILES标注的长上下文科学文档组成,分别包含104亿和115亿个token。随后,模型使用从原始训练语料库通过算法衍生的SMILES-文本对合成数据集进行对比学习。该设计使模型接触到富含SMILES的科学文献,从而实现双语义理解。我们在涵盖分子性质预测和科学语言理解的一系列下游任务上评估CheMatE。结果表明,将我们定制整理的数据集与该序列训练策略相结合,可产生稳健且高度可迁移的表示。通过在单一基于文本的框架内有效统一结构与上下文信号,CheMatE在专用化学模型和通用语言模型基准上均取得了有竞争力的性能。

英文摘要

Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry. However, domain-adaptive pre-training often causes models to overfit to chemical syntax, catastrophically forgetting their foundational semantic capabilities. To address this challenge, we introduce CheMatE, a chemistry-oriented embedding model that jointly captures molecular structure and domain-specific natural language within the same representation space. Built on a ModernBERT backbone, CheMatE learns bi-semantic representations through a two-stage training procedure: continued masked language modeling (MLM) followed by a Matryoshka contrastive learning stage via Multiple Negative Ranking Loss (MNRL). First, we train the model using MLM on a novel, large-scale corpus of SMILES-annotated, long-context scientific documents that were constructed and curated from FineWeb and ChemPile (comprising 10.4B and 11.5B tokens, respectively). Subsequently, the model undergoes contrastive learning using a synthetic dataset of SMILES-text pairs algorithmically derived from our original training corpus. This design exposes the model to SMILES-enriched scientific literature, enabling bi-semantic understanding. We evaluate CheMatE across a range of downstream tasks covering molecular property prediction and scientific language understanding. Our results demonstrate that coupling our custom-curated datasets with this sequential training strategy yields robust, highly transferable representations. By effectively unifying structural and contextual signals within a single text-based framework, CheMatE achieves competitive performance across both specialized chemistry models and general-purpose language model baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑