arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于高性能计算环境中可解释机器日志分析的可扩展模式挖掘工作流程

A Scalable Pattern Mining Workflow for Interpretable Machine Log Analysis in High-Performance Computing Environments

Shilpika Shilpika, Bethany Lusch, Eric Pershey, Carlo Graziani, Venkatram Vishwanath, Michael E. Papka

arXiv 2607.19143首次发表:更新:

AI 中文总结

研究高性能计算环境中机器日志分析问题,结合先进模式匹配与挖掘技术,利用系统层次结构和消息优先级提取模式序列,关联聚类错误序列与作业日志,经案例研究证明该方法可释放日志数据潜力,助力HPC系统优化。

AI 中文摘要

现代高性能计算(HPC)环境中的超级计算机每天都会生成大量日志数据,揭示这些复杂系统的详细信息和性能指标。HPC日志的规模和异构性,尤其是文本数据,给传统分析技术带来了重大挑战。因此,分析这些日志时需要更复杂的工作流程来提取模式,以发现可能指示系统故障的潜在模式和异常,并帮助预测未来的故障和效率低下情况。我们的日志分析工作流程研究了应用于HPC日志分析的先进模式匹配和挖掘技术的组合。通过系统地识别日志消息中的频繁日志模式和模式序列,并将它们存储在有限状态自动机(如Aho-Corasick自动机)中,我们的工作流程能够自动检测频繁错误和故障事件。为了提取这些模式和序列,我们利用系统层次结构和消息优先级的信息。然后,我们将识别出的错误序列与作业日志进行关联和聚类,揭示具有相似或不同错误特征的应用程序组。这种方法产生的见解为改进提供了依据,并指导实时监控工作。我们的研究表明,模式挖掘对于通过实现实时分析并为更具弹性、可扩展的HPC系统做出贡献来释放日志数据的全部潜力至关重要。我们通过汇总统计和对百亿亿次超级计算机的案例研究证明了我们方法的有效性。

英文摘要

Modern supercomputers housed in High Performance Computing (HPC) environments generate massive volumes of log data daily, revealing intricate information and performance metrics about these complex systems. The sheer size and heterogeneous nature of HPC logs, especially text data, pose significant challenges for traditional analytical techniques. Consequently, more complex workflows are necessary for pattern extraction when analyzing these logs, enabling the discovery of underlying patterns and anomalies that may indicate system faults and help predict future failures and inefficiencies. Our log analysis workflow investigates a combination of advanced pattern-matching and mining techniques applied to HPC log analysis. By systematically identifying frequent log patterns and pattern sequences in log messages and storing them in a finite-state automaton, such as the Aho-Corasick automaton, our workflow enables automated detection of frequent errors and fault events. To extract these patterns and sequences, we leverage information about system hierarchy and message priority. We then correlate and cluster the identified error sequences with job logs, revealing groups of applications with similar or dissimilar error signatures. This approach yields insights that inform improvements and guide real-time monitoring efforts. Our research establishes that pattern mining is vital for unlocking the full potential of log data by enabling real-time analysis and contributing to more resilient, scalable HPC systems. We demonstrate the effectiveness of our approach through summary statistics and a case study on an exascale-class system supercomputer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑