TwinGridShield:面向LLM网格智能体动作的后果感知运行时授权
TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions
浏览论文内容
中文总结 AI 辅助
TwinGridShield是与模型无关的LLM网格智能体动作运行时授权层,通过确定性网络孪生体评估动作,在匹配模型实验中实现0次不安全发布,模型不匹配下鲁棒性存在一定风险。
中文摘要 AI 辅助
基于大语言模型(LLM)的能源管理工具可将自然语言上下文转换为结构化电网指令,但语法有效不代表物理上可接受。本文提出TwinGridShield,一种与模型无关的运行时授权层,在每个拟议动作发布前通过确定性网络孪生体对其进行评估。原型系统检查连通性、分支潮流、发电机及切负荷不变量,并将每个决策记录在哈希链日志中。受控IEEE 14节点研究采用直流潮流和实验分配的分支额定值,评估单步开关、再调度及切负荷动作。在匹配模型实验中,配置为以概率p=0.84选择不安全动作的随机提议源,在500次攻击条件试验中产生421个不安全提议,实际比率为84.2%,该值表征所配置的替代模型,而非LLM提示注入敏感性的经验测量。TwinGridShield在这500次试验中产生0次不安全发布。由于动作标记和授权使用相同的直流模型、系统状态、分支额定值及编码约束,该结果验证了实现与其编码授权谓词的一致性,而非模型误差下的安全性。因此,主要鲁棒性评估引入模型不匹配:在每节点负荷测量误差有界为±20%时,不安全接受率达5.63%;当实际分支额定值比建模值低20%时,该比率为30.09%。
英文摘要
Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a model-independent runtime authorization layer that evaluates each proposed action in a deterministic network twin before release. The prototype checks connectivity, branch-flow, generator, and load-shedding invariants and records each decision in a hash-chained log. A controlled IEEE 14-bus study evaluates single-step switching, redispatch, and load-shedding actions using DC power flow and experimentally assigned branch ratings. In the matched-model experiment, a stochastic proposal source configured to select an unsafe action with probability p=0.84 produced 421 unsafe proposals in 500 attacked-condition trials, a realized rate of 84.2%. This value characterizes the configured surrogate and is not an empirical measurement of LLM prompt-injection susceptibility. TwinGridShield produced 0 unsafe releases in those 500 trials. Because action labeling and authorization used the same DC model, system state, branch ratings, and encoded constraints, this result verifies conformance of the implementation to its encoded authorization predicate rather than safety under model error. The principal robustness evaluation therefore introduces model mismatch. Unsafe acceptance reached 5.63% under bounded +20% and -20% per-bus load-measurement error and 30.09% when actual branch ratings were 20% below modeled ratings.
发表机构
- West Virginia University(西弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。