arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VeraGrid-Agent:用于电网边缘分布式最优潮流的工具增强大语言模型

VeraGrid-Agent: Tool-Augmented LLMs for Distribution Optimal Power Flow at the Grid Edge

Shivanshu Tripathi, Hamed Mohsenian-Rad, Maziar Raissi

arXiv 2607.25155首次发表:更新:

AI 中文总结

研究针对电网边缘分布式最优潮流问题,提出工具增强大语言模型VeraGrid-Agent,通过自主编写模拟器输入、执行求解器并读取输出作答。引入VeraGrid-MCQ-150测试评估,相比无工具推理,该模型显著提升准确率,且剩余错误源于推理而非求解器执行。

AI 中文摘要

语言模型在解决广泛任务方面取得了显著成功。然而,回答有关潮流的复杂科学问题通常需要解决分布式最优潮流(D-OPF)问题,语言推理往往给出错误答案。本文提出VeraGrid-Agent,这是一种工具增强的大语言模型,其能自主编写模拟器输入,执行开源VeraGrid求解器并读取输出后作答。为评估性能,引入了一项有150个选择题的VeraGrid-MCQ-150测试。在无工具推理和带有模拟器访问的代理两种模式下评估性能,无工具时模型准确率为42.7%至49.3%,使用VeraGrid-Agent后准确率提高到97.3%至100.0%。还进行了失败模式分析,表明剩余错误源于多步推理中的错误解释,而非模拟器执行失败。

英文摘要

Language models have demonstrated remarkable success in solving a wide range of tasks. However, answering complex scientific questions about the power flow often requires solving the distribution optimal power flow (D-OPF) problem. These questions call for numerical solvers and simulators, as linguistic reasoning from parametric knowledge often gives incorrect answers. In this work, we present VeraGrid-Agent, a tool-augmented LLM that autonomously writes the simulator input, executes the open-source VeraGrid solver, and reads the solver output before answering. To evaluate performance, we introduce VeraGrid-MCQ-150, a set of deterministic, expert template driven, $150$ multiple-choice questions. We evaluate the performance under two regimes: (i) no-tool reasoning and (ii) agent (LLM with simulator access). Without tools, every model performs with an accuracy of $42.7\%$--$49.3\%$. However, with VeraGrid-Agent, accuracy increases to $97.3\%$--$100.0\%$. We also do a failure-mode analysis to show that the few remaining errors arise from wrong interpretations during multi-step reasoning, rather than any failure in the simulators execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑