发表机构
Laboratory of Scientific Computing and Visualization, Federal University of Alagoas; Computing Institute, Federal University of Alagoas(阿拉戈斯联邦大学科学计算与可视化实验室; 阿拉戈斯联邦大学计算学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种部署在油井开放世界异常检测OWL流程下游的LLM智能体层,借助Qwen3.5-397B-A17B模型,可解释上游决策、标记分歧、命名新异常,缩小OWL流程的可解释性差距。
AI 中文摘要
用于油井异常检测的开放世界学习(OWL)流程近期被证实可结合基于自编码器的检测、多分类分类及基于马氏距离的新异常检测方法,且基于公开的3W数据集实现。这些流程能回答“发生了什么”,但无法解释“模型为何会这样认为”或“操作人员接下来应做什么”,也无法为其发现的新异常簇赋予人类可读的名称。本文评估了一个部署在OWL流程下游的大语言模型(LLM)智能体层,该层被设计为已发布上游方法的“辅助工具”而非替代方案。该智能体使用通过NVIDIA NIM服务的Qwen3.5-397B-A17B混合专家模型,接收结构化传感器指标及上游分类或新异常断言,并返回自然语言的理由、按置信度排序的评论以及检测到的新异常的统一名称。在针对3W数据集中989个真实井段的三项研究中,该智能体在全部9个类别上实现了35.1%的top-1/63.9%的top-3(95%置信区间[56.9,70.4])分类准确率,在7个被探测类别上实现了71.7%的top-2验证准确率[64.8,77.6],精确率为0.91[0.84,0.95],在7个隐藏类别中的5个上实现了89.7%的新异常检测准确率[87.0,91.9]及稳定的簇命名。该智能体并非独立分类器,其作用为:(1)当传感器证据支持上游决策时确认该决策;(2)以操作人员可审计的、基于传感器的语言对决策进行解释;(3)当上游标签不合理时标记分歧;(4)为新异常命名,使聚类后的未标记事件能以统一的人类可读标签传递给工程师。该研究的目标是缩小目前阻碍OWL流程在实际场景中部署的可解释性差距。
英文摘要
Open-World Learning (OWL) pipelines for oil well anomaly detection have recently been shown to combine autoencoder-based detection, multiclass classification, and Mahalanobis-based novelty detection on the public 3W dataset. These pipelines answer \textit{what happened}, but they do not explain \textit{why the model believes it} or \textit{what the operator should do next}, and they do not put a human-readable name on the novelty clusters they discover. This paper evaluates a Large Language Model (LLM) agent layer placed downstream of the OWL pipeline, designed as a \textbf{companion} to the published upstream methods rather than a replacement. Using the Qwen3.5-397B-A17B Mixture-of-Experts model served via NVIDIA NIM, the agent receives structured sensor metrics and upstream classification or novelty assertions, and returns natural-language justifications, confidence-ranked critiques, and consolidated names for detected novelties. Across three studies spanning 989 real well-file segments from the 3W dataset, the agent achieved $35.1\%$ top-1 / $63.9\%$ top-3 (95\% CI [56.9, 70.4]) classification on all nine classes, $71.7\%$ top-2 validation [64.8, 77.6] with precision $0.91$ [0.84, 0.95] across 7 probed classes, and $89.7\%$ novelty detection [87.0, 91.9] with stable cluster naming on 5 of 7 hidden classes. The agent is not a standalone classifier. Its role is to: (1) confirm upstream decisions when sensor evidence supports them, (2) justify decisions in sensor-grounded language operators can audit, (3) flag disagreement when upstream labels are implausible, and (4) name novelties so that clustered unlabeled events arrive at the engineer with a consolidated human-readable label. The goal is to close the explainability gap that currently blocks deployment of OWL pipelines in operational settings.