发表机构
The University of Texas at Arlington; Brookhaven National Laboratory; Southern Methodist University(德克萨斯大学阿灵顿分校; 布鲁克海文国家实验室; 南卫理公会大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态大语言模型用于电网诊断时答案准确性与证据使用的问题,提出通用框架进行任务条件忠实性审计,通过比较多种依赖并设计校正和重新审计机制,经案例研究验证了框架检测、诊断及纠正相关失败的能力。
AI 中文摘要
多模态大语言模型可结合拓扑、测量和事件文本进行电网诊断,但答案准确性不能表明使用了与任务相关的证据。本文提出了一个通用框架来进行任务条件忠实性审计。它比较了自我报告的依赖、干预产生的行为依赖和预先注册的工程重要性。该框架首先记录特定任务的证据要求,并将其与自我报告的依赖以及在受控模态消融下的行为变化进行比较。为解决检测到的差异,设计了一种证据门控校正和重新审计机制,在证据约束下重新生成失败的响应,并独立地重新消融它们以验证改进的基础且不损失性能。案例研究在IEEE 39和118总线场景中评估了三个不同规模的大语言模型。这些结果验证了该框架检测、诊断和纠正任务条件忠实性失败的能力。
英文摘要
Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with self-reported reliance and behavioral changes under controlled modality ablations. To resolve detected discrepancies, we design an evidence-gated correction and re-audit mechanism that regenerates failed responses under evidence constraints and independently re-ablates them to verify improved grounding without performance loss. Case studies evaluate three differently scaled LLMs on IEEE 39- and 118-bus scenarios. These results validate the framework ability to detect, diagnose, and correct task-conditional faithfulness failures.