AI 中文总结
研究针对MOF数据库CIF输入错误影响下游结果及人工检查的问题,提出MOF-Sleuth,通过强化引导CIF审核代理的两个模块,利用奖励引导强化学习将工具测量转化为监督,提升检测、归因及解释质量,在多基准测试中性能领先。
AI 中文摘要
大型金属有机框架(MOF)数据库通过晶体学信息文件(CIF)支持模拟、筛选和机器学习。这些输入中的细微化学和结构错误会影响下游结果并阻碍人工检查。计算化学中语言模型的进步为细粒度诊断提供了超越预测筛选的途径,并提供有证据支持的解释。然而,存在两个挑战:细粒度归因有限,MOF特定验证器和机器学习模型扩大了检测范围,但提供固定检查、准备分数或粗略标签,而非有证据支持的解释;CIF推理不可靠,直接的语言模型审核成本高且不可靠,因为化学证据在原子位点记录中是隐含的,需要进行几何、连接性、占有率和电荷计算。这两个问题都源于化学证据与语言模型解释之间的弱耦合。我们引入了MOF-Sleuth,这是一个强化引导的CIF审核代理,有两个模块:确定性法医实验室和Sleuth推理引擎。实验室得出组成、几何、连接性、占有率、配位和电荷证据,Sleuth利用这些证据产生有证据支持的解释、错误类型和二元决策。奖励引导的强化学习将工具测量转化为化学解释级别的监督,不仅奖励最终答案,还奖励引用的化学证据和有证据支持的诊断。我们引入了化学基础诊断(Chem-GD),这是一种评估正确诊断是否由基于事实的、相关的CIF衍生证据解释的指标。在四个基准测试中,MOF-Sleuth在基于语言模型的方法和MOF特定的机器学习方法中建立了领先的性能,在检测、归因和有根据的解释质量方面都有提升。
英文摘要
Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural errors in these inputs can compromise downstream results and hinder manual inspection. LLM advances in computational chemistry offer paths beyond predictive screening toward fine-grained diagnosis with evidence-grounded explanations. However, two challenges remain: (i) limited fine-grained attribution: MOF-specific validators and machine-learning models scale detection but provide fixed checks, readiness scores, or coarse labels rather than evidence-grounded explanations; and (ii) unreliable CIF reasoning: direct LLM auditing is costly and unreliable because chemical evidence is implicit across atom-site records and requires geometric, connectivity, occupancy, and charge calculations. Both stem from weak coupling between chemical evidence and language-model explanation. We introduce MOF-Sleuth, a reinforcement-guided CIF auditing agent with two modules: a deterministic Forensic Lab and a Sleuth reasoning engine. The Lab derives composition, geometry, connectivity, occupancy, coordination, and charge evidence, and Sleuth uses this evidence to produce an evidence-grounded explanation, error types, and a binary decision. Reward-guided reinforcement learning (RL) turns tool measurements into chemical explanation-level supervision, rewarding not only the final answer but also cited chemical evidence and evidence-supported diagnoses. We introduce Chemically Grounded Diagnosis (Chem-GD), a metric that assesses whether a correct diagnosis is explained by factual, relevant CIF-derived evidence. Across four benchmarks, MOF-Sleuth establishes state-of-the-art performance among LLM-based approaches and MOF-specific machine-learning methods, demonstrating gains in detection, attribution, and grounded explanation quality.