发表机构
ETH Zürich; Politecnico di Milano; Harvard University(苏黎世联邦理工学院; 米兰理工大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出RestoreBench基准,在两种电网的92个潮流案例上评估多类LLM及三种架构的智能体恢复潮流收敛的能力,为电力系统智能体AI开发提供可复现基础。
AI 中文摘要
大型语言模型(LLM)智能体正通过工具使用、中间结果解读及迭代规划,日益实现多步骤工程工作流的自动化。诊断并解决不收敛的潮流案例是一个有前景但大多未被探索的应用,因为它需要工程判断、实验以及在受限动作空间内的决策。我们引入一个基准,在多个LLM和三种架构(聊天机器人、单智能体、多智能体系统)上评估这些能力。该评估覆盖两个电网,每个电网含46个案例,每个案例需一次或多次纠正动作以恢复收敛。该基准定义了模拟环境、观测与动作空间及评估指标,为开发用于电力系统规划与运行的智能AI系统提供了可复现的基础,代码可在该https URL获取。
英文摘要
Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use, interpretation of intermediate results, and iterative planning. Diagnosing and resolving non-convergent power flow cases is a promising yet largely unexplored application, as it requires engineering judgment, experimentation, and decision-making within constrained action spaces. We introduce a benchmark that evaluates these capabilities across multiple LLMs and three architectures: \emph{chatbot}, \emph{single agent}, and \emph{multi-agent} systems. The evaluation covers two power grids and 46 cases per grid, each requiring one or more corrective actions to restore convergence. The benchmark defines the simulation environment, observation and action spaces, and evaluation metrics, providing a reproducible foundation for developing agentic AI systems for power system planning and operation. The code is available at https://github.com/Mansutti081/RestoreBench