arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ClouDens:用于大规模云系统监控的操作上下文感知异常检测

ClouDens: Operational Context-Aware Anomaly Detection for Large-scale Cloud System Monitoring

Thu T. H. Doan, Mohammad Saiful Islam, Andriy Miranskyy, Ngoc-Thanh Nguyen, Rogardt Heldal, Patrizio Pelliccione

arXiv 2607.18127首次发表:更新:

发表机构

Department of Computer Science, Gran Sasso Science Institute; University of Bergen; Department of Computer Science, Toronto Metropolitan University; Department of Computer Science, Electrical Engineering and Mathematical Sciences, Western Norway University of Applied Sciences(计算机科学系,格兰萨索科学研究所; 卑尔根大学; 计算机科学系,多伦多 Metropolitan 大学; 计算机科学、电气工程和数学科学系,西挪威应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大规模云系统监控中异常检测难题,提出ClouDens框架。它利用遥测日志模式中的操作上下文属性,将高维日志划分子集,构建上下文感知图,采用时空图神经网络进行异常检测,在相关数据集上表现良好,为云系统监控异常检测提供指导。

AI 中文摘要

随着云计算基础设施在规模和复杂性上的快速增长,大规模云系统(LCS)的网络监控变得越来越具有挑战性,需要自动化和可靠的异常检测来维持服务可用性。现代LCS不断从分布式云服务生成遥测日志,产生捕获系统操作的高维多元时间序列。由于维度极高、分布式组件之间的复杂依赖关系以及间歇性活跃服务导致的严重稀疏性,在此背景下检测异常很困难。考虑到这些挑战,我们首先对来自IBM云控制台平台的遥测日志进行了实证研究,然后提出了ClouDens,这是一个针对LCS监控量身定制的异常检测框架,它利用遥测日志模式中编码的操作上下文属性来提高检测准确性和异常的早期识别。ClouDens将高维遥测日志划分为领域引导的子集,构建一个上下文感知图来建模操作服务依赖关系,并采用时空图神经网络进行基于预测的异常检测。我们在最近发布的IBM云遥测数据集上对ClouDens进行了评估,并为设计用于LCS监控的可靠异常检测解决方案提供了实际见解。结果表明,ClouDens在基于计数的遥测特征方面实现了更高的NAB分数,表明比基于GRU的模型具有更准确、更早的异常检测和更广泛的覆盖范围。我们的研究进一步揭示,遥测特征子集、操作上下文建模、评分策略和稀疏性插补都对检测性能有重大影响,为设计和公平基准测试LCS监控的异常检测方法提供了实际指导。

英文摘要

With the rapid growth of cloud computing infrastructures in scale and complexity, network monitoring for Large-scale Cloud Systems (LCSs) has become increasingly challenging, requiring automated and reliable anomaly detection to maintain service availability. Modern LCSs continuously generate telemetry logs from distributed cloud services, producing high-dimensional multivariate time series that capture system operations. Detecting anomalies in this context is difficult due to extreme dimensionality, complex dependencies among distributed components, and severe sparsity from intermittently active services. Taking these challenges into account, we first conduct an empirical study on telemetry logs from the IBM Cloud Console platform, and then propose ClouDens, an anomaly detection framework tailored to LCS monitoring that leverages operational-context attributes encoded in the telemetry log schema to improve detection accuracy and early identification of anomalies. ClouDens partitions high-dimensional telemetry logs into domain-guided subsets, constructs a context-aware graph modeling operational service dependencies, and employs Spatio-Temporal Graph Neural Networks for forecasting-based anomaly detection. We evaluate ClouDens on the recently released IBM Cloud Telemetry Dataset and provide practical insights into designing reliable anomaly detection solutions for LCS monitoring. Results show ClouDens achieves higher NAB scores in count-based telemetry features, indicating more accurate, earlier anomaly detection with broader coverage than a GRU-based model. Our study further reveals that telemetry feature subsets, operational-context modeling, scoring strategies, and sparsity imputation all substantially influence detection performance, offering practical guidance for designing and fairly benchmarking anomaly detection approaches for LCS monitoring.

Comments16 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑