深度生成晶体结构预测:基准研究及原型依赖性的受控测试
Deep Generative Crystal Structure Prediction: A Benchmark Study and a Controlled Test of Prototype Dependence
查看机构详情
- University of South Carolina(南卡罗来纳大学)
- University of North Georgia(北佐治亚大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究基准测试12种深度生成晶体结构预测模型,发现其性能高度依赖模板原型,真正从头预测能力有限,扩大残余预测能力是核心挑战。
中文摘要 AI 辅助
深度生成模型被广泛报道能够实现从头晶体结构预测(CSP),但其能力尚未与基于模板的方法进行一致的衡量。我们评估了12个具有代表性的生成式CSP模型,涵盖潜变量、扩散、流匹配、自回归和流形随机游走架构,在TCSP 2.0上对180个测试结构以及一个由46个结构组成的泄漏控制子集进行了测试。所有方法均使用相同的结构匹配、对称性和一致性标准。模板检索是最强的单一方法,达到68.3%的top-1成功率;对称感知的EquiCSP(66.4%)和Uni-3DAR(62.9%)构成下一梯队。然而,与TCSP 2.0的比较表明,生成模型正确预测的大多数结构也能通过模板替换正确预测。因此,生成模型能够唯一到达的结构集合很小,限制了其在现有原型库之外发现结构的实际优势。为了测试这种性能的来源,我们从训练集中移除了整个化学计量原型家族,并重新训练了最强的生成模型。在四个家族中,准确率下降了50-78%,表明性能在很大程度上依赖于原型。少数结构在移除其原型家族后仍然存活,证明了存在真实但有限的、不依赖检索的预测能力。因此,当前的生成式CSP模型在很大程度上充当了隐式的、边缘模糊的原型库,而非真正的从头预测器。扩大这种残余能力,而不仅仅是提高总体匹配率,是核心的开放问题。
英文摘要
Deep generative models are widely reported to enable de novo crystal structure prediction (CSP), but their capability has not been measured consistently against template-based methods. We evaluate 12 representative generative CSP models, spanning latent-variable, diffusion, flow-matching, autoregressive, and manifold random-walk architectures, against TCSP 2.0 on 180 test structures and a leakage-controlled subset of 46. All methods use identical structure-matching, symmetry, and consensus criteria. Template retrieval is the strongest single method, reaching 68.3% top-1 success; symmetry-aware EquiCSP (66.4%) and Uni-3DAR (62.9%) form the next tier. However, comparison with TCSP 2.0 shows that most structures correctly predicted by generative models are also correctly predicted by template substitution. Thus, the set of structures uniquely reachable by generation is small, limiting its practical advantage for discovering structures outside existing prototype libraries. To test the source of this performance, we removed entire stoichiometric prototype families from the training set and retrained the strongest generative model. Accuracy declined by 50-78% across four families, establishing that performance is substantially prototype-dependent. A small minority of structures survived removal of their prototype family, demonstrating a real but limited retrieval-independent predictive capacity. Present generative CSP models therefore function largely as implicit, softer-edged prototype libraries rather than genuinely de novo predictors. Enlarging this residual capacity, rather than aggregate match rate alone, is the central open problem.