AI 中文总结
Zatom-2是在OMol25和OMat24数据集上预训练的原子生成模型,采用多尺度Transformer与条件流匹配,提升了分子生成保真度,增强了低数据下的蛋白质生成能力。
AI 中文摘要
统一原子建模有望通过连接数据丰富的化学领域与数据稀缺的生物领域,加速化学、材料科学和生物学的发现。然而,现有的原子建模生成方法仍高度专业化(如针对化学或生物学),或未同时利用大量有机(分子)和无机(材料)数据进行通用预训练。为此,我们提出Zatom-2,这是一种原子生成模型,在来自OMol25和OMat24电子结构数据集的约500万个结构上进行预训练。Zatom-2具有多尺度Transformer架构,结合了支持力条件化的条件流匹配,以及生成、结构预测、分子和材料能量与力预测等基础预训练任务。实验表明,Zatom-2比Zatom-1实现了更好的分子分布保真度,在现有分子和材料生成基准上表现出色;它还能控制低、高力 regime下的样本生成,通过联合生成-预测预训练和迁移学习,在低数据场景下增强蛋白质生成,在长度外推场景中,蛋白质骨架可设计性从无预训练时的67.8%提升至在2000个蛋白质域上微调后的74.8%。
英文摘要
Unified atomistic modeling has the potential to accelerate discovery in chemistry, materials science, and biology by bridging data-rich chemical domains and data-scarce biological contexts. However, existing generative approaches to atomistic modeling remain highly specialized to scientific disciplines (chemistry vs. biology) or do not leverage both high-volume organic (molecule) and inorganic (material) data for general-purpose pretraining. To this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets. Zatom-2 features a multiscale Transformer architecture coupled with conditional flow matching that supports force conditioning and foundational pretraining tasks such as generation, structure prediction, and prediction of molecular and material energies and forces. Empirically, Zatom-2 achieves better molecular distribution fidelity than Zatom-1 and achieves strong performance on existing molecule and material generation benchmarks. Zatom-2 demonstrates the ability to control sample generation across low- and high-force regimes, and enhances protein generation in a low-data setting through joint generative-predictive pretraining and transfer learning, increasing protein backbone designability in a length extrapolation setting from 67.8% without pretraining to 74.8% after finetuning on 2,000 protein domains.