发表机构
Center for Physical Sciences and Technology (FTMC); Institute of Biotechnology, Life Sciences Center, Vilnius University(立陶宛物理与技术中心; 维尔纽斯大学生命科学中心生物技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估通用机器学习原子间势能(MLIPs)能否自动化晶体学开放数据库(COD)的结构物理合理性验证,发现其为有用但依赖模型的筛选工具,结合多个模型更优。
AI 中文摘要
晶体学开放数据库(COD)包含超过50万个实验测定的晶体结构。晶体学数据的验证分为三个层级,第三层级评估结构的物理合理性。鉴于晶体学数据库的规模和新结构的提交速率,该层级通常依赖经验启发式方法或专家判断,速度慢且易出错。本文评估通用机器学习原子间势能(MLIPs)是否可实现该层级验证的自动化。在522086条COD条目中,317457条(占60.8%)通过了数据质量过滤,该过滤移除了无序结构、化学式不匹配及未建模溶剂。这些结构用M3GNet进行弛豫,其中314153条(占99.0%)收敛且晶胞体积无大幅变化。对于大多数条目,弛豫体积与沉积体积几乎一致,因此收敛性和体积变化可作为识别无法确认沉积结构的弛豫的简单标准。对327条结构的人工检查显示,没有单一描述符阈值(能量、最大力或原子位移)能区分有效与无效结构。在524条选定结构上对M3GNet、CHGNet、PET-MAD和PET-OAM的比较显示出模型特定的失效情况:M3GNet会使环戊二烯基及其他π配体发生畸变,M3GNet和CHGNet均会使噻吩和噻唑环发生畸变,而PET模型能保留这些结构基序。所有模型均检测到缺失的氢原子,至少部分模型检测到虚假氢原子、错误分配的原子类型以及两种此前未报道的坐标误差,且没有单一模型能检测到所有误差类型。因此,MLIP弛豫是一种有用但依赖模型的筛选工具,结合多个模型比依赖单一模型更可取。
英文摘要
The Crystallography Open Database (COD) contains over half a million experimentally determined crystal structures. Validation of crystallographic data proceeds at three hierarchical levels, the third of which assesses the physical plausibility of structures. This level has typically relied on empirical heuristics or expert judgment, which are slow and error-prone given the size of crystallographic databases and the rate at which new structures are deposited. Here we assess whether universal machine-learning interatomic potentials (MLIPs) can automate this level of validation. Of the 522 086 COD entries, 317 457 (60.8\%) passed a data-quality filter removing disordered structures, formula mismatches and unmodelled solvent. These structures were relaxed with M3GNet, and 314 153 (99.0\%) converged without a large change in cell volume. For most entries the relaxed and deposited volumes are nearly identical. Convergence and volume change thus provide simple criteria for identifying relaxations that do not confirm the deposited structure. A manual inspection of 327 structures showed that no single descriptor threshold (energy, largest force, or atomic displacement) separates valid from invalid structures. A comparison of M3GNet, CHGNet, PET-MAD and PET-OAM on 524 selected structures revealed model-specific failures: M3GNet distorts cyclopentadienyl and other $π$ ligands, and both M3GNet and CHGNet distort thiophene and thiazole rings, whereas the PET models preserve these motifs. All models detected missing hydrogen atoms, and at least some detected spurious hydrogens, incorrectly assigned atom types, and two previously unreported coordinate errors. No single model detected all error types. MLIP relaxation is therefore a useful but model-dependent screening tool, and combining several models is preferable to relying on any single one.
Comments33 pages, 12 figures, 5 tables