arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34473cs.SE

从噪声遥测到可操作警告:工业集群中的GPU故障预测

From Noisy Telemetry to Actionable Warnings: GPU Failure Prediction in Industrial Clusters

  • Nankai University(南开大学)
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

Yongqian Sun, Run Zhu, Wenwei Gu, Mengyao Li, Shenglin Zhang, Guanjin Wang, Yang Zhang, Xin Wu, Linlin Han, Feng Wang, Xiaozhou Liu, Yu Zhang

AI总结:

针对生产GPU集群中噪声遥测导致的故障预测难题,提出故障特定框架Falcon,结合缺失感知特征与事件策略,在测试集上取得最高F1并实现数小时提前预警。

AI中文摘要:

GPU集群是AI服务的关键基础设施,但在生产环境中实现准确且可操作的GPU故障预测仍是一个难题。我们研究了字节跳动GPU集群中与工单关联的遥测数据,并识别出三个障碍:工作负载混杂的遥测、异构的故障前兆,以及窗口级预测与可操作警报之间的差距。这些发现促使我们提出Falcon,一个故障特定的警告框架,结合了缺失感知的时间特征和对等相对特征、故障特定的学习器选择,以及基于阈值、持续和冷却的事件策略。在测试集上,Falcon在四个基线中取得了最高的F1分数,并在性能最佳的故障类型上达到70.6%的F1。检测到的案例提供了17.34-35.57小时的中位提前时间。我们进一步报告了一次生产部署,其中Falcon被校准为高精度警报以反映误报成本。综合这些结果,表明故障特定的建模改善了从噪声生产GPU遥测中的早期警告。

英文摘要:

GPU clusters are critical infrastructure for AI services, but accurate and actionable GPU failure prediction remains a problem in production settings. We study ticket-linked telemetry from a ByteDance GPU cluster and identify three obstacles: workload-confounded telemetry, heterogeneous fault precursors, and the gap between window-level predictions and actionable alerts. These findings motivate Falcon, a fault-specific warning framework combining missingness-aware temporal and peer-relative features, fault-specific learner selection, and an event policy based on thresholding, persistence, and cooldown. On the test set, Falcon achieves the highest F1 among four baselines and reaches 70.6% F1 on the best-performing fault type. Detected cases provide median lead times of 17.34-35.57 hours. We further report a production deployment, where Falcon is calibrated toward high-precision alerts to reflect false-positive costs. Together, these results show that fault-specific modeling improves early warning from noisy production GPU telemetry.

补充信息

↑