arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于Volve油田的接地良好异常检测:构建标签、基线及双头模型

Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model

Gospel Bassey, Samuel Bassey, Vincent Fakiyesi

arXiv 2608.05685首次发表:更新:

AI 中文总结

该研究针对Volve油田数据构建经工程文档验证的异常标签,测试其可学习性,提出双头模型,公开相关资源,为实际生产场景的异常检测提供基准与模型支持。

AI 中文摘要

大多数用于机器状态监测的公开基准数据集来自试验台,这些试验台会人为引入故障,且所有事件都是已知的。而实际生产油田很少能提供这种条件,它们给出的是未附带故障日志的传感器历史数据,这正是异常检测方法必须自行生成标签,且可能悄无声息地混入不合理假设的场景。我们使用Equinor发布的公开Volve油田数据,重点关注这类数据集通常忽略的两点:其一,我们构建的异常标签并非仅基于数据中的模式,而是对照油田自身工程文档中记载的可能物理故障进行验证,且公布每个标签背后的推理依据;其二,我们通过无监督基线模型和小型双头模型测试这些构建的标签是否可学习,该双头模型用于标记事件发生与否及事件类型,这一思路源自早期金属部件缺陷检测的研究。结果是可信的:从未见过这些标签的无监督检测器仍能定位到我们规则标记的相同区域,表明这些标签并非随意生成;紧凑的监督模型在从未见过的油井上能较好地恢复事件存在性和类型,但时间定位仅较为粗略。我们报告了有效的方法、无效的内容以及其间的所有假设。该数据集、接地标签、每个标签的来源、基线分数、训练好的模型及代码均以CC-BY-NC-SA 4.0许可公开。

英文摘要

Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known. Real production fields rarely offer that. They give you sensor histories with no fault log attached, which is exactly the situation where an anomaly-detection method has to invent its own labels, and where quiet assumptions can slip in unnoticed. We work with the open Volve field data released by Equinor and take two things seriously that such datasets usually skip. First, we build anomaly labels that are not just patterns in the numbers but are checked against what the field's own engineering documents say can physically go wrong, and we release the reasoning behind every label. Second, we test whether those constructed labels are learnable at all, using both an unsupervised baseline and a small dual-head model that marks when an event happens and what kind it is, an idea we carry over from earlier work on defect detection in metal parts. The results are honest. An unsupervised detector that never sees the labels still lands on the same regions our rules flagged, which tells us the labels are not arbitrary. A compact supervised model recovers event presence and event type well across wells it has never seen, and locates events in time only roughly. We report what worked, what did not, and every assumption in between. The dataset, grounded labels, per-label provenance, baseline scores, trained model, and code are released publicly under CC-BY-NC-SA 4.0.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑