发表机构
Lawrence Berkeley National Laboratory; University of California, Berkeley; Science and Technology Facilities Council(劳伦斯伯克利国家实验室; 加州大学伯克利分校; 科学与技术设施委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出uMOF数据库、基准及两款通用MLIP,在MOF的动力学敏感属性预测上性能优于现有基准,误差降低80%以上。
AI 中文摘要
基础机器学习原子间势(MLIP)能以远低于从头算的计算成本实现接近从头算的精度,但其在金属有机框架(MOF)领域的潜力仍未得到充分发挥,原因在于大晶胞使得生成第一性原理训练数据成本高昂、微调模型稀缺、基于实验的基准则更为稀少。为此,我们推出uMOF,包含三部分内容以填补这一空白。第一,我们发布了迄今为止最大、最准确的MOF密度泛函理论数据集,采用r$^2$SCAN-D4理论级别计算,涵盖19950种独特框架、79种元素的85524个构型,包括空结构与负载气体的结构、几何优化、物态方程及有限温度分子动力学模拟。第二,我们发布了经文献挖掘的基准数据集,包含3986个验证后的属性值(其中3146个为实验值),这些值通过一个七阶段、带检查点的多轮大语言模型流程从626篇论文中提取,并与超过650个晶体学信息文件关联。第三,我们发布了两个针对MOF的通用MLIP,即uMOF-MH和uMOF-POLAR,它们是基于两种架构不同的MACE基础模型在uMOF数据集上微调得到的。在近平衡的“Tier-1”属性(体积模量、声子衍生热容)上,uMOF模型的表现与现有基础模型及微调基准相当;在更难、对动力学敏感的属性(如通过Widom插入法得到的气体吸附焓和吸附等温线)上,uMOF模型优于我们测试的所有基准,包括在大三个数量级的数据集上训练的MOF专用气体捕获模型,其误差降低了80%以上,达到实验不确定度范围内。我们将这一优势归因于训练数据的物理多样性,以及分子动力学模拟中1.7%的小比例数据对MLIP稳定性具有决定性作用的理论级别。
英文摘要
Foundation machine learning interatomic potentials (MLIPs) deliver near-ab-initio accuracy at a fraction of the computational cost, yet their promise for Metal-organic Frameworks (MOFs) remains largely unrealized as large unit cells make first-principles training data expensive to generate, fine-tuned models are scarce, and experimentally grounded benchmarks are scarcer still. We introduce uMOF, a three-part contribution addressing this gap. First, we release the largest and most accurate density functional theory dataset for MOFs to date, computed at the r$^2$SCAN-D4 level of theory across 85524 configurations spanning 19950 unique frameworks and 79 elements, covering empty and gas-loaded structures, geometry optimizations, equations of state, and finite-temperature molecular dynamics. Second, we release a literature-mined benchmark of 3986 verified property values (3146 experimental) extracted from 626 papers by a seven-stage, checkpointed multi-pass large language model pipeline, linked to more than 650 crystallographic information files. Third, we release two universal MLIPs for MOFs, uMOF-MH and uMOF-POLAR, fine-tuned from two architecturally distinct MACE foundation models on the uMOF dataset. On near-equilibrium, ``Tier-1'' properties (bulk modulus, phonon-derived heat capacity) the uMOF models perform comparably to existing foundation and fine-tuned baselines. On harder, dynamics-sensitive properties like gas adsorption enthalpies via Widom insertion and adsorption isotherms, the uMOF models outperform every baseline we test, including MOF-specialized gas-capture models trained on datasets up to three orders of magnitude larger, cutting error by more than 80% to within experimental uncertainty. We trace this advantage to the physical diversity of the training data and to level of theory where a small (1.7%) fraction of MD simulations is decisive for MLIP stability.
Comments19 pages, 5 figures, 7 tables