PowerBench:电力系统中智能体检索与推理的基准
PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems
浏览论文内容
中文总结 AI 辅助
针对电力系统智能体评估缺乏真实数据与链式依赖的问题,提出PowerBench基准,含生成框架和合成数据集,测试前沿LLM在受限条件下的检索推理,最佳准确率仅74.2%,为可靠部署提供洞见。
中文摘要 AI 辅助
大语言模型(LLM)智能体为工业领域的自动化分析提供了新的机遇。然而,对此类智能体(例如在电力系统场景中)的严格评估仍受到阻碍:真实运行数据具有保密性,而现有的公开资源未能充分捕捉链式依赖和异构证据。为弥补这一空白,我们提出了PowerBench,包括(1)一个通过公共依赖链生成相互关联的异构运行数据的生成框架,以及(2)由该框架生成的合成数据集。该数据集覆盖100种设备类型中的761台设备,包含跨越两年的1335万条小时级遥测记录和24,939份运行文档。基于该数据集,我们构建了三个任务族中的300个问题,用于评估前沿LLM在受限工具调用和时间预算下,完成需要自主证据检索和跨相互关联及异构数据推理的分析任务的能力。结果表明,所评估的前沿LLM在这些任务上仍面临挑战:最佳模型仅达到74.2%的联合准确率。我们的轨迹分析进一步揭示,模型性能在证据发现、内容检索、工具使用、证据推理和答案提交方面存在差异。这些发现为评估LLM智能体并指导其在工业中的可靠部署提供了详细见解。框架、数据集和基准任务可在以下https URL获取。
英文摘要
Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs' ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at https://github.com/open-compass/PowerBench.
发表机构
- Tongji University(同济大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Big Data Center of State Grid Corporation of China(国家电网有限公司大数据中心)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- Carnegie Mellon University(卡内基梅隆大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。