物理应从哪里进入分子晶体生成器?
Where Should Physics Enter a Molecular Crystal Generator?
浏览论文内容
中文总结 AI 辅助
本文提出全原子晶体流映射生成模型CrystAF,结合UMA势系统比较物理在训练后与推理时介入的效果,发现训练后学习物理偏好可提升分子有效性与堆积质量且不增采样成本,推理时校正互补修复残余冲突,该策略跨架构可迁移。
中文摘要 AI 辅助
生成模型使分子晶体结构预测变得快速,但其样本仍存在几何和堆积违例。物理可以在训练期间、训练后或推理时引入,然而这些选择很少在生成器和物理信号固定的情况下进行比较。我们提出了CrystAF,一种全原子晶体流映射生成模型,并将其与UMA原子间势结合,系统研究物理应从哪里进入。训练后学习将物理偏好直接融入CrystAF,提高了分子有效性和晶体堆积,同时保持采样不变:物理在训练期间一次性支付,而不是在部署时重复支付。相比之下,UMA弛豫能有效修复局部冲突,但使生成速度降低6至26倍,而从弛豫目标学习则收益甚微。这些路线是互补而非竞争的。物理信息训练后首先将生成分布转向更物理合理的结构,随后廉价的推理时校正进一步消除冲突并恢复生成器无法表示的立体化学。重要的是,相同的训练后策略也改进了多步全原子Clari-M和刚体MolCrystalFlow生成器,展示了跨架构和表示的迁移性。综合来看,我们的结果提出了一个简单原则:将可复用的物理对齐学习到生成器中,并将推理时物理保留给那些更适合校正而非学习的残余约束。
英文摘要
Generative models make molecular crystal structure prediction fast, but their samples still exhibit geometric and packing violations. Physics can be introduced during training, post-training, or inference, yet these choices are rarely compared with the generator and physical signal held fixed. We introduce CrystAF, an all-atom crystal flow-map generation model, and use it with the UMA interatomic potential to systematically study where physics should enter. Post-training learns physical preferences directly into CrystAF, improving molecular validity and crystal packing while leaving sampling unchanged: physics is paid for once during training rather than repeatedly at deployment. In contrast, UMA relaxation is effective at repairing local clashes but makes generation 6--26$\times$ slower, while learning from relaxed targets provides little benefit. These routes are complementary rather than competing. Physics-informed post-training first shifts the generated distribution toward more physically reasonable structures, after which inexpensive inference-time corrections further remove clashes and restore stereochemistry that the generator cannot represent. Importantly, the same post-training strategy also improves the multi-step all-atom Clari-M and rigid-body MolCrystalFlow generators, demonstrating transfer across architectures and representations. Together, our results suggest a simple principle: learn reusable physical alignment into the generator, and reserve inference-time physics for residual constraints that are better corrected than learned.
发表机构
- Khoury College of Computer Science(胡里计算机学院)
- School of Pharmacy(药学院)
- Northeastern University(东北大学)
- University of Pittsburgh(匹兹堡大学)
机构由 AI 辅助整理,请以论文原文为准。