arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不要预测,进行优先级排序:重新思考GPU可靠性评估

Don't Predict, Prioritize: Rethinking GPU Reliability Assessment

Difeng Ma, Changhua Pei, Yuanwei Lu, Quan Zhou, Zexin Wang, Yibo Zhu, Daxin Jiang, Dan Pei, Jingjing Li, Gaogang Xie

arXiv 2607.15115首次发表:更新:

AI 中文总结

研究针对GPU可靠性评估难题,传统预测方法受限,提出HeaRank排序学习框架按相对故障风险对节点排序,在生产集群评估中表现优异,能有效捕获未来故障,凸显风险感知调度和主动资源管理在GPU集群中的重要性。

AI 中文摘要

图形处理单元(GPU)的可靠性是现代大规模人工智能基础设施的关键瓶颈,单个节点故障会扰乱同步训练作业并造成重大经济损失。虽然预测性维护在其他硬件领域广泛应用,但准确预测GPU故障的确切时间具有内在局限性。通过对生产集群遥测数据的深入分析,发现主要GPU故障具有很强的随机性和低信噪比,传统基于时间的预测无效。因此提出范式转变,不再预测故障绝对时间,而是按相对故障风险对节点进行排序。提出HeaRank(健康排名),一种利用稳定历史故障模式计算GPU节点全局风险排名的排序学习框架。在拥有数千个GPU的生产规模集群上评估,HeaRank的AUC为0.83,显著优于启发式基线和现有排名算法。在线部署中,HeaRank在前5%排名节点中成功捕获64%的未来故障,而现有生产系统仅为21%。这些结果表明,在绝对故障预测固有受限的环境中,相对风险排名可作为有力替代方案。强调了现代GPU集群中风险感知调度和主动资源管理的重要性。

英文摘要

The reliability of Graphics Processing Units (GPUs) is a criticalbottleneck for modern large-scale AI infrastructure, where a sin-gle node failure can disrupt synchronous training jobs and causesignificant financial losses. While predictive maintenance is widelyused in other hardware domains, we demonstrate that accuratelypredicting the exact timing of GPU failures is inherently difficult.Through an in-depth analysis of telemetry data from a productioncluster, we find that major GPU failures, including Double Bit Er-rors (DBEs) and GPU Lost events, exhibit strong stochasticity andlow signal-to-noise ratios in time-series telemetry, which makesconventional time-based prediction ineffective. This insight motivates a paradigm shift: instead of attempting topredict the absolute timing of a failure, we propose a more robustapproach focused on ranking nodes by their relative failure risk. Wepropose HeaRank (Health Rank), a Learning-to-Rank (LTR) frame-work that leverages stable historical failure patterns to computea global risk ranking of GPU nodes. Evaluated on a production-scale cluster with thousands of GPUs, HeaRank achieves an AUCof 0.83, significantly outperforming both heuristic baselines andstate-of-the-art ranking algorithms. In online deployment, HeaRanksuccessfully captures 64% of future failures within the top 5% ofranked nodes, compared to only 21% by the incumbent productionsystem. These results suggest that relative risk ranking can serveas a robust alternative in environments where absolute failure pre-diction is inherently limited. Our work highlights the importanceof risk-aware scheduling and proactive resource management inmodern GPU clusters.

CommentsAccepted at ACM SIGKDD 2026; 13 pages, 13 figures

Journal refProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026)

DOI:10.1145/3770855.3818373

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑