arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MLIP Detective:超越基准分数的机器学习原子间势主动失效模式发现

MLIP Detective: Active Failure Mode Discovery Beyond Benchmark Scores for Machine-Learning Interatomic Potentials

Ryuhei Okuno, Nontawat Charoenphakdee, Kaoru Hisama, Yuta Tsuboi

arXiv 2609.08399首次发表:更新:

发表机构

Preferred Networks(Preferred Networks)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MLIP Detective智能体框架,通过物理信息搜索主动发现机器学习原子间势的隐藏失效模式,并在MACE-MPA-0中识别出吸附物-表面体系能量异常及其训练数据来源。

AI 中文摘要

通用机器学习原子间势(u-MLIPs)旨在泛化到各种不同的构型。基准测试能够实现可复现的评估,但可能无法暴露其预定义范围之外的失效情况。在此,我们表明,物理信息搜索可以补充基于基准的评估,通过揭示隐藏的失效模式。我们引入了MLIP Detective,一个用于主动失效模式发现的智能体框架。从基准证据出发,MLIP Detective生成可证伪的、基于物理的失效假设,通过廉价模拟进行筛选,并仅将最可疑的案例升级给人类专家,同时附上提议的验证协议。在没有针对特定问题的提示的情况下,MLIP Detective识别并表征了MACE-MPA-0中的一个系统性异常:该模型预测某些涉及含O或F吸附物的松弛吸附物-表面体系比其对应的分离碎片能量更高。通过跨模型比较,MLIP Detective进一步推断出该异常可能的训练数据来源,这与最近的报告一致。

英文摘要

Universal machine-learning interatomic potentials (u-MLIPs) aim to generalize across diverse configurations. Benchmarks enable reproducible evaluation but may not expose failures outside their predefined scope. Here, we show that physics-informed search can complement benchmark-based evaluation by uncovering hidden failure modes. We introduce MLIP Detective, an agentic framework for active failure mode discovery. Starting from benchmark evidence, MLIP Detective generates falsifiable, physics-informed failure hypotheses, screens them with inexpensive simulations, and escalates only the most suspicious cases to human experts together with proposed verification protocols. Without issue-specific prompting, MLIP Detective identified and characterized a systematic anomaly in MACE-MPA-0: the model predicted some relaxed adsorbate-surface systems involving O- or F-containing adsorbates to be higher in energy than their corresponding separated fragments. Using cross-model comparisons, MLIP Detective further inferred a likely training-data origin for the anomaly, consistent with recent reports.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑