arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22476cs.SEcs.AI

面向输电系统运营商控制室决策支持的治理感知大语言模型编排智能体数字孪生

A Governance-Aware Large Language Model Orchestrated Agentic Digital Twin for Transmission System Operator Control Room Decision Support

Costas Mylonas, Magda Foti, Emmanouel Varvarigos

首次发表
浏览论文内容

中文总结 AI 辅助

针对输电控制室决策,提出治理感知智能体数字孪生,由大语言模型仅选白名单工具并受不可绕过治理层约束,在118任务基准上实现93.7%成功率且规则零违反。

中文摘要 AI 辅助

输电系统运营商面临可再生能源并网、惯量降低和安全裕度收紧带来的日益增长的复杂性。大语言模型提供自然语言决策支持,但其幻觉、不受控制的工具使用和薄弱的可追溯性与控制室要求相冲突。本文提出了一种面向输电网控制室的治理感知智能体数字孪生。大语言模型仅能选择和参数化白名单内的分析工具,且每个提议的动作都必须经过模型无法绕过的治理层。该层在每次运行中强制执行四条规则:仅白名单工具可执行;每次运行不得超过其步骤预算;任何具有副作用的动作未经操作员明确批准不得执行;答案中的每个数字均由该层根据后端结果渲染,并附有其单位、变量和时间。这些规则在希腊输电网数字孪生上发布的118任务基准的每次运行中均通过持久化审计轨迹进行检查,该基准涵盖分析、仿真、多步骤工作流以及十二类对抗性输入。在主要模型的590次运行中,工具选择达到96.5%,任务成功率达93.7%,且四条规则无一例外全部成立。在另一项针对四个大语言模型、每次重复三次、总计1416次运行的独立研究中,规则在每次运行中同样成立,近似95%置信下限为99.8%。移除该层后,同一模型在45次需要批准的运行中全部未经授权执行,且其答案中仅有39.2%的数字由后端支持。强制执行成本为每次请求12至16毫秒。

英文摘要

Transmission system operators face rising complexity from renewable integration, reduced inertia, and tighter security margins. Large language models offer natural-language decision support, but their hallucinations, uncontrolled tool use, and weak traceability conflict with control room requirements. This paper presents a governance-aware agentic digital twin for transmission grid control rooms. The large language model only selects and parameterizes whitelisted analysis tools, and every proposed action passes through a governance layer that the model cannot bypass. The layer enforces four rules on every run. Only whitelisted tools execute. No run exceeds its step budget. No action with side effects executes without explicit operator approval. Every number in an answer is rendered by the layer from backend results with its unit, variable, and time. The rules are checked on a persistent audit trail for every run of a released 118-task benchmark, which covers analytics, simulation, multi-step workflows, and twelve families of adversarial inputs on a digital twin of the Greek transmission network. Across 590 runs of the primary model, tool selection reaches 96.5% and task success 93.7%, and all four rules hold without exception. In a separate three-repetition study across four large language models, 1416 runs in total, the rules again hold on every run, with an approximate 95% lower bound of 99.8%. Removing the layer makes the same model execute all 45 approval-requiring runs without authorization and leaves only 39.2% of its answers with backend-supported numbers. Enforcement costs 12 to 16 milliseconds per request.

发表机构

  • UBITECH(UBITECH公司)
  • National Technical University of Athens(雅典国立技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑