评估大型语言模型在强迫停运风险预测中的应用:优势及与机器学习的对比
Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning
查看机构详情
- Temple University(天普大学)
- Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究对比零样本LLMs与监督分类器预测配电网强迫停运风险的表现,发现监督模型在宏F1和精确率上更优,LLMs在可操作推理等方面有互补优势,二者结合或为最佳实践。
中文摘要 AI 辅助
本研究在零样本框架下,考察大型语言模型(LLMs)预测配电网与天气相关强迫停运风险的能力,无需带标注的训练数据。该问题被构造成二元严重程度分类任务,覆盖三个预测时域(3小时、6小时、12小时),使用美国得克萨斯州中部某公用事业服务区六年的停运记录和高分辨率天气数据。在两种输入配置下,将四个零样本LLMs与两个监督分类器进行基准测试:一种使用当前天气观测数据,另一种使用天气预报数据。结果显示,监督模型在宏F1值和精确率上优于LLMs,而较新的LLM代际取得了具有竞争力的分数。除准确性外,LLMs在可操作推理和地理可扩展性方面具备互补优势,表明将其与监督模型结合可能是最佳实践。
英文摘要
This study examines the ability of large language models (LLMs) to predict the risk of weather-related forced outages in the distribution grid in a zero-shot framework, without labeled training data. The problem is formulated as a binary severity classification task across three forecast horizons (3h, 6h, 12h), using six years of outage records and high-resolution weather data for a utility service area in central Texas. Four zero-shot LLMs are benchmarked against two supervised classifiers across two input configurations: one using current weather observations and the other using weather forecast data. Results show that supervised models outperform LLMs on macro-F1 and precision, while newer LLM generations achieve competitive scores. Beyond accuracy, LLMs offer complementary strengths in actionable reasoning and geographic scalability, suggesting that combining them with supervised models may be the best practice.