arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ElectrolyteMD-Bench:AI智能体在电解质溶剂化区间中进行分子动力学研究的成效如何?

ElectrolyteMD-Bench: How Well Do AI Agents Conduct Molecular Dynamics Studies across Electrolyte Solvation Regimes?

Shukai Wu, Shuo Niu, Zhaoming Xu, Yan Luo, Wentao Lin

arXiv 2609.31743首次发表:更新:

AI 中文总结

提出ElectrolyteMD-Bench基准,评估AI智能体在三种电解质溶剂化区间中自主完成MD研究的能力,发现所有配置均未能满足全部科学要求,揭示了工作流执行与可靠研究完成之间的差距。

AI 中文摘要

AI智能体正在推进分子动力学(MD)研究的自动化,有望加速锂离子电池电解质的发现与设计。然而,它们能否自主完成科学上可靠的电解质MD研究仍未得到充分评估。我们提出了ElectrolyteMD-Bench,该基准将稀溶液、高浓度溶液和局部高浓度电解质统一在一个跨不同溶剂化区间的比较研究中。该基准将局部配位、离子缔合、单粒子动力学和集体输运连接在一个连续的研究过程中。智能体在冻结的科学模型下自主执行系统构建、平衡、生产采样、性质分析和解释,同时决定是否纠正错误、继续采样或停止。我们沿四个维度对每项研究进行审计:模拟有效性、分析有效性、证据完整性和结果校准。在八种模型-框架配置中,所有24个系统运行均产生了有效的生产轨迹,但八项研究均未满足所有科学要求。失败包括分析缺失、物理定义或数值实现错误、统计支持不足,以及完成声明与保留证据不一致。这些结果揭示了成功的MD工作流执行与可靠科学研究完成之间的差距。因此,ElectrolyteMD-Bench提供了一个基础测试,并定义了自主电解质模拟必须跨越的关键能力门槛,以推进以计算为先的材料探索。

英文摘要

AI agents are advancing the automation of molecular dynamics (MD) research, with the potential to accelerate discovery and design of lithium-battery electrolytes. However, whether they can autonomously complete scientifically reliable electrolyte MD studies remains insufficiently evaluated. We introduce ElectrolyteMD-Bench, which unifies dilute, high-concentration, and localized high-concentration electrolytes in a comparative study across distinct solvation regimes. The benchmark connects local coordination, ionic association, single-particle dynamics, and collective transport within one continuous research process. Agents autonomously perform system construction, equilibration, production sampling, property analysis, and interpretation under a frozen scientific model, while deciding whether to correct errors, continue sampling, or stop. We audit each study along four dimensions: simulation validity, analysis validity, evidence integrity, and outcome calibration. Across eight model--harness configurations, all 24 system runs produced valid production trajectories, yet none of the eight studies satisfied all scientific requirements. Failures included missing analyses, incorrect physical definitions or numerical implementations, insufficient statistical support, and completion claims inconsistent with retained evidence. These results expose a gap between successful MD workflow execution and reliable scientific study completion. ElectrolyteMD-Bench therefore provides a foundational test and defines a key capability threshold that autonomous electrolyte simulation must cross to advance toward computation-first materials exploration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑