arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13357physics.chem-phcond-mat.mtrl-sciphysics.comp-ph

自主LLM智能体能否执行多参考量子化学计算?

Can Autonomous LLM Agents Execute Multireference Quantum Chemistry Calculations?

  • Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

Victor Chang Lee, James M. Rondinelli

AI总结:

本研究测试自主LLM智能体执行多参考量子化学计算的能力,通过结构化决策阶梯提升活性空间选择与态识别,显著提高覆盖率并降低误差,证明其可复现复杂工作流。

AI中文摘要:

多参考电子结构计算的自动化仍然困难,因为关键的工作流决策,包括活性空间选择、态平均、收敛恢复和态识别,传统上依赖于专家判断。在此,我们研究自主大语言模型(LLM)智能体是否能在无需人工干预的情况下执行这些任务。该智能体使用基于文献的类比或明确记录的化学推理来选择活性空间,生成并提交ORCA计算,分析输出,并将所有决策记录在可审计的推理日志中。针对QUESTDB中的558个垂直跃迁能(VTE)进行基准测试表明,无引导的基线智能体实现了24.9%的覆盖率,平均绝对误差(MAE)为0.373 eV。引入结构化决策阶梯后,覆盖率提高到44.1%,同时MAE降至0.339 eV。最大的改进出现在双重和里德伯激发中,表明专家知情的程序性指导显著改善了活性空间构建和态识别。当提供完整的工作流信息时,智能体成功复现了已发表的QUEST计算,MAE仅为23 meV,并在七次尝试内解决了75%的目标构型。这些结果表明,当代LLM智能体能够自主执行和复现复杂的多参考量子化学工作流,同时强调了结构化推理框架对于实现可靠的高通量和高保真电子结构计算的重要性。

英文摘要:

Multireference electronic-structure calculations remain difficult to automate because critical workflow decisions, including active-space selection, state averaging, convergence recovery, and state identification, traditionally rely on expert judgment. Here, we investigate whether an autonomous large language model (LLM) agent can perform these tasks without human intervention. The agent selects active spaces using literature-grounded analogies or explicitly documented chemical reasoning, generates and submits ORCA calculations, analyzes outputs, and records all decisions in an auditable reasoning log. Benchmarking against 558 vertical transition energies (VTEs) from QUESTDB shows that an unguided baseline agent achieves 24.9% coverage with a mean absolute error (MAE) of 0.373 eV. Introducing a structured decision ladder increases coverage to 44.1% while reducing the MAE to 0.339 eV. The largest gains are observed for double and Rydberg excitations, demonstrating that expert-informed procedural guidance substantially improves active-space construction and state identification. When provided with complete workflow information, the agent successfully reproduces published QUEST calculations with an MAE of only 23 meV, resolving 75% of target configurations within seven attempts. These results demonstrate that contemporary LLM agents can autonomously execute and reproduce complex multireference quantum-chemical workflows, while highlighting the importance of structured reasoning frameworks for achieving reliable high-throughput and high-fidelity electronic-structure calculations.

补充信息

↑