arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有化学意义的文本化方法可通过大语言模型实现金属有机框架的可解释性验证

Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models

Guobin Zhao, Xiao-Yan Li

arXiv 2608.11283首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出将晶体学信息转换为具有化学意义的文本的方法,微调LLM(基于mof2text)可实现MOF结构的可解释性验证,性能与图模型相当,还能生成错误诊断依据。

AI 中文摘要

可用于计算的金属有机框架(MOF)数据库对高通量筛选至关重要,但许多已报道的晶体结构仍存在化学不合理性或无序性,损害了模拟的保真度。现有的验证方法能够识别出不可用于计算的结构,但通常依赖启发式规则、授权要求,或可解释性有限。本文展示,当晶体学信息被转换为具有化学意义的文本时,大语言模型(LLM)可作为MOF结构的可解释验证器。通过对9种描述符进行基准测试,我们发现基于LLM的验证的成功并不单独取决于结构信息的数量,而是取决于局部配位、框架连接性和化学背景是否被组织成语言可学习的表示形式。使用专用描述符(mof2text)微调的LLM在识别不合理MOF方面达到了与基于图的模型相当的性能。重要的是,这些模型超越了黑箱分类,能够生成潜在错误来源的诊断依据,包括异常键合、连接性和电荷状态,以及带注释数据集的错误类别预测。本研究确立了化学信息文本化是将LLM从通用文本模型转变为用于整理MOF数据库的实用可解释工具的关键步骤。

英文摘要

Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity. Existing validation approaches can identify non-computation-ready structures, but they often rely on heuristic rules, license requirement, or offer limited interpretability. Here, we show that large language models (LLMs) can serve as interpretable validators of MOF structures when crystallographic information is transformed into chemically meaningful text. By benchmarking nine descriptors, we find that successful LLM-based validation depends not on the amount of structural information alone, but on whether local coordination, framework connectivity, and chemical context are organized into a linguistically learnable representation. Fine-tuned LLMs using specialized descriptors (mof2text) achieve performance comparable to graph-based models in identifying unreasonable MOFs. Importantly, these models extend beyond black-box classification by generating diagnostic rationales for likely error sources, including abnormal bonding, connectivity, and charge states, as well as error-category predictions for annotated datasets. This work establishes chemically informed textualization as the key step that transforms LLMs from generic text models into practical and explainable tools for curating MOF databases.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑