arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于化学中生成模型和基础模型的高阶分子文法

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang, Jure Leskovec, Tolga Birdal

arXiv 2610.02186首次发表:更新:

发表机构

Queen Mary University of London; Stanford University; Imperial College London; Yonsei University(伦敦玛丽女王大学; 斯坦福大学; 帝国理工学院; 延世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出高阶文法表示(HGR),将分子解析为规则序列以高效编码高阶拓扑,并构建富环基准RingDiv,在生成和表示学习任务中均取得领先性能。

AI 中文摘要

分子学习模型深受其底层表示的影响。然而,标准的序列和图形式体系难以显式编码高阶拓扑结构,如环系统和重复出现的基序。现有的高阶表示可以直接捕获这些结构,但通常计算要求高且难以解码为有效的分子。在此,我们引入高阶文法表示(HGR),这是一个有原则的、拓扑感知的框架,它将分子提升为组合复形,并在上下文无关的高阶文法下将每个复形解析为一系列紧凑的产生式规则。通过将高阶拓扑序列化为规则序列,HGR使这些结构直接兼容标准序列模型,避免了显式高阶编码的计算开销,同时保留了拓扑表达能力。为了减少对简单环系统的基准偏差,我们构建了RingDiv,一个包含118万个分子的富环基准,包括精选的RingDiv300k子集,并引入环多样性指数(RDI)来量化环系统覆盖率。在分子生成中,基于HGR的模型独特地结合了构造上的100%有效性和领先的分布对齐,在所有五个生成基准上FCD排名第一。在表示学习中,HGR-FM在两种迁移协议下在七个MoleculeNet基准上实现了最高的平均AUC,在探测和全微调下分别比最强基线提高了8.3和3.3个AUC点。总的来说,这些结果确立了HGR作为分子生成和可迁移表示学习的高效高阶表示。

英文摘要

Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑