arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04037cs.DC

BitIR:面向异构GPU应用韧性分析的跨架构故障注入

BitIR: Cross-Architecture Fault Injection for Resilience Analysis of Heterogeneous GPU Applications

Maisy Dunlavy, Michael Papka, Zhiling Lan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出跨架构故障注入框架BitIR,在LLVM IR层注入单比特故障,对比NVIDIA、Intel、AMD GPU上的韧性表现,发现故障行为因后端而异,强调需后端与工作负载感知的故障缓解。

中文摘要 AI 辅助

现代基于GPU的HPC系统依赖于异构供应商技术栈,然而韧性研究大多局限于单一架构,导致故障在不同GPU后端之间的行为尚不明确。我们提出了BitIR,一种跨架构故障注入框架,在LLVM IR层面注入确定性的单比特故障,确保在后端降级之前产生语义等价的扰动,从而支持在NVIDIA、Intel和AMD GPU之间进行直接的跨供应商比较。我们在三台生产级超级计算机上评估了BitIR——Polaris(ALCF,NVIDIA A100)、Aurora(ALCF,Intel GPU Max 1550)和Frontier(OLCF,AMD Instinct MI250X)——它们代表了当前领导级GPU架构的全谱系。使用具有代表性的异构基准测试,我们在所有三个系统上开展了大规模注入活动,并将结果分类为掩蔽结果、静默数据损坏(SDC)和故障。我们的结果表明,相同的故障在不同供应商之间产生显著不同的行为:Intel最常掩蔽故障,但在故障逃脱掩蔽时表现出更多的挂起;NVIDIA暴露更多可检测的硬故障;而AMD则根据基准测试和故障位置在SDC主导和故障主导行为之间交替。这些发现表明韧性并非后端不变的,强调了需要后端感知和工作负载感知的故障缓解措施。

英文摘要

Modern GPU-based HPC systems rely on heterogeneous vendor stacks, yet resilience studies are largely limited to single architectures, leaving it unclear how faults behave across different GPU backends. We present \emph{BitIR}, a cross-architecture fault injection framework that injects deterministic single-bit faults at the LLVM IR level, ensuring semantically equivalent perturbations prior to backend lowering and enabling direct cross-vendor comparison across NVIDIA, Intel, and AMD GPUs. We evaluate BitIR on three production supercomputers -- Polaris (ALCF, NVIDIA A100), Aurora (ALCF, Intel GPU Max 1550), and Frontier (OLCF, AMD Instinct MI250X) -- representing the full spectrum of current leadership-class GPU architectures. Using representative heterogeneous benchmarks, we conduct large-scale injection campaigns across all three systems and classify outcomes into masked results, silent data corruptions (SDCs), and failures. Our results show that identical faults produce markedly different behaviors across vendors: Intel most often masks faults but exhibits more hangs when faults escape masking, NVIDIA exposes more detectable hard failures, and AMD alternates between SDC-dominant and failure-dominant behavior depending on the benchmark and fault site. These findings demonstrate that resilience is not backend-invariant, underscoring the need for backend-aware and workload-aware fault mitigation.

发表机构

  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)
  • Argonne National Laboratory(阿贡国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑