arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Monroe:用于上下文概率推理的分子基础模型

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Blazej Banaszewski, Andrew W. Fitzgibbon

arXiv 2608.18982首次发表:更新:

发表机构

Graphcore; Max Planck Institute of Biochemistry(格弗科(Graphcore); 马克斯·普朗克生物化学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出新的分子基础模型Monroe,通过多方面创新提升性能,在多个基准测试中表现优异,且其下游适应策略可推广至其他模型。

AI 中文摘要

生物测定活性预测常受限于数据不足,因为药物发现数据集依赖耗时且昂贵的湿实验室实验来生成和评估数据。这一挑战推动了近期对分子基础模型(Molecular Foundation Models, MFMs)的研究,这类模型旨在将通用化学知识编码为分子表示,以便在数据受限场景中实现良好泛化。本文提出了Monroe,一种新的分子基础模型,它在现有最优模型基础上有多项创新:规模扩大,可在PM6量子化学数据集的超过8100万个分子上进行预训练;改进了立体化学的图表示;改进了训练损失,包括构象去噪和嵌入去相关;改进了多任务学习;以及使用先验数据拟合模型TabPFN进行下游上下文预测。我们的评估采用了原则性的成对比较框架,用于衡量具有统计显著性的性能差异。在已建立的Polaris基准测试中,Monroe与现有分子基础模型相当或更优;在旨在评估分子发现实用性的活性崖基准测试中,它比先前方法实现了显著改进。最后,消融实验和迁移实验表明,基于PFN的下游预测器也大幅改进了两个领先的现有模型MiniMol和CheMeleon,产生了新的最优变体MiniMol_PFN和CheMeleon_PFN,这表明我们的下游适应策略可推广到Monroe之外。源代码位于此http URL。

英文摘要

Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. This challenge has inspired recent research into molecular foundation models (MFMs), which aim to encode general-purpose chemical knowledge into molecular representations that generalize well in data-constrained scenarios. This paper presents Monroe, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 81 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction. Our evaluations use a principled pairwise comparison framework that measures statistically significant performance differences. Across established Polaris benchmarks, Monroe matches or exceeds existing MFMs, while on activity cliff benchmarks, designed to assess utility for molecular discovery, it achieves significant improvements over prior methods. Finally, ablation and transfer experiments show that PFN-based downstream predictors also substantially improve two leading existing models, MiniMol and CheMeleon, yielding new state-of-the-art variants we call MiniMol_PFN and CheMeleon_PFN, suggesting that our downstream adaptation strategy generalizes beyond Monroe. Source code is at github.com/blazejba/monroe.

CommentsPreprint; Open source weights and code

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑