arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37158cs.LGcond-mat.mtrl-sci

GLASS:基于槽集合解码的全局潜在聚合,用于可扩展的全原子晶体生成

GLASS: Global Latent Aggregation with Slot-based Set Decoding for Scalable All-Atom Crystal Generation

Hendrik Kraß, Seyed Mohamad Moosavi, Mathias Niepert

首次发表
浏览论文内容

中文总结 AI 辅助

GLASS通过置换不变的全局潜在空间和槽解码器消除粒子对应问题,实现可扩展的全原子晶体生成,在MP20和QMOF上达到高有效性。

中文摘要 AI 辅助

晶体的生成模型能够发现新结构,但将全原子生成扩展到金属有机框架等更大系统仍然具有挑战性。我们将这一困难与粒子空间生成的对应问题联系起来。即使在单个固定目标集上,在独立和最优传输耦合下,随着集合大小和密度的增加,无索引置换等变粒子流也需要更多的训练才能实现可靠的生成。为解决这一挑战,我们引入了GLASS——全局潜在聚合与基于槽的集合解码,它在置换不变的全局潜在空间中编码结构,并通过流匹配学习其分布。一个学习槽解码器并行构建所有原子,从生成传输中消除了原子级对应关系。在MP20上,GLASS与粒子空间模型具有竞争力,且流训练在每个结构尺寸下都能达到训练数据的有效性。在QMOF子集上,GLASS在不依赖构建块、拓扑或组成的情况下生成每晶胞最多150个原子的MOF,并接近训练数据的结构有效性。在两个数据集上,流训练揭示了有效性与新颖性之间的权衡,且MOF的新颖性仍受限于可用数据上的自编码器泛化能力。这些结果表明,将对应分配与生成传输分离为更大原子系统的高有效性生成提供了一条简单途径。

英文摘要

Generative models for crystals enable the discovery of novel structures, but scaling all-atom generation to larger systems such as metal--organic frameworks remains challenging. We connect this difficulty to the correspondence problem of particle-space generation. Even on a single fixed target set, index-free permutation-equivariant particle flows require substantially more training for reliable generation as set size and density increase, under both independent and optimal-transport couplings. To resolve this challenge, we introduce GLASS---Global Latent Aggregation with Slot-based Set Decoding, which encodes structures in a permutation-invariant global latent space and learns their distribution via flow matching. A learned-slot decoder constructs all atoms in parallel, removing atom-wise correspondence from generative transport. On MP20, GLASS is competitive with particle-space models, and flow training can reach the validity of the training data at every structure size. On a QMOF subset, GLASS generates MOFs with up to 150 atoms per unit cell without conditioning on building blocks, topology, or composition, and approaches the structural validity of the training data. On both datasets, flow training exposes a validity--novelty tradeoff, and MOF novelty remains limited by autoencoder generalization on the available data. These results show that separating correspondence assignment from generative transport provides a simple route toward high-validity generation of larger atomistic systems.

发表机构

  • Institute for Artificial Intelligence(人工智能研究所)
  • University of Stuttgart(斯图加特大学)
  • University of Toronto(多伦多大学)
  • Vector Institute(向量研究所)
  • NEC Labs Europe(NEC欧洲实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑