发表机构
University of California, Davis; Argonne National Laboratory; Rensselaer Polytechnic Institute(加州大学戴维斯分校; 阿贡国家实验室; 伦斯勒理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究基于Theta超级计算机7年生产日志,评估经典统计与深度学习模型对HPC硬件错误的预测效能,明确不同时间结构错误的预测可行性,为相关分析提供实证指导。
AI 中文摘要
高性能计算(HPC)系统中的硬件错误日志可提供异常行为的早期信号,但采用现代预测方法有效预测这些错误仍存在挑战。本研究探究将时间序列预测应用于HPC硬件错误动态的边界,使用Theta超级计算机7年的生产日志,评估经典统计模型与深度学习模型的预测效能。结果表明,预测有效性高度依赖错误序列的时间结构:规律出现且结构稳定的错误可被准确建模,尤其是具备时间特征的LSTM和Transformer架构;而稀疏且突发主导的错误仍难以预测。本研究未提出可部署的故障预测框架,而是提供了关于预测何时有效的实证指导,并指出了提升HPC硬件错误分析预测精度的潜在方向。
英文摘要
Hardware error logs in high-performance computing (HPC) systems provide early signals of abnormal behavior, yet there remain challenges in effectively forecasting these errors using modern predictive methods. This work investigates the boundaries of applying time series forecasting to HPC hardware error dynamics. We use seven years of production logs from the Theta supercomputer to evaluate the predictive efficacy of classical statistical and deep learning models. Our results show that forecasting effectiveness depends strongly on the temporal structure of the error series: regularly occurring and structurally stable errors can be modeled accurately, particularly by LSTM and Transformer architectures with temporal features, while sparse and burst-dominated errors remain difficult to predict. Rather than proposing a deployment-ready failure prediction framework, this study provides empirical guidance on when forecasting is effective and highlights potential directions for improving forecasting accuracy in HPC hardware error analysis.
CommentsAccepted at the 7th International Workshop on Monitoring, Observability, and Operational Data Analytics (MODA 2026), held in conjunction with ISC High Performance 2026