发表机构
Advanced Micro Devices, Inc.(Advanced Micro Devices 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在多种工作负载和浮点格式中表征DNN推理内存的位位置故障敏感性,发现位敏感性转变,提出不等错误保护方法,得出数据类型下限和工作负载感知层级,驱动UEP编解码器及双分区架构,减少ECC面积、降低读取能量。
AI 中文摘要
我们在16种工作负载(涵盖基于Transformer的模型和无注意力的卷积神经网络)以及三种浮点格式中,对机器学习推理中的每位位置故障敏感性进行了表征。我们的核心实证发现是位敏感性的急剧转变:在确定性单比特压力测试下,翻转任何最低有效分数比特直至特定数据类型的阈值Xsafe,任务指标的下降幅度小于1%。敏感性通过较高分数比特上升,并在指数-尾数边界处激增,此时单比特翻转会导致灾难性崩溃。由于低阶比特基本无关紧要而高阶和指数比特至关重要,均匀的SECDED保护(以12.5%的存储开销平等保护每个比特)过于保守。我们得出了每种数据类型的Xsafe下限(FP16:6,BF16:4,FP32:15)和工作负载感知层级,为弹性模型类别扩大未受保护区域,在无需重新训练的情况下将ECC节省提高到37.5 - 62.5%。文本条件扩散模型决定了保守下限;视觉编码器、自然语言理解模型和弹性语言模型容忍更宽的旁路区域。这些下限和层级驱动了一种具有每个缓存行数据类型标签的不等错误保护(UEP)编解码器以及用于机器学习加速器的双分区SRAM架构。通过870多次故障注入运行的验证证实,在连续2比特和3比特翻转下,选择性保护依然成立。该编解码器相对于均匀的SECDED将ECC面积减少了27.8%;非关键分区的双电压操作将BF16的总读取能量降低了约17%,双分区宏面积开销约为4%。
英文摘要
We characterize per-bit-position fault sensitivity in ML inference across 16 workloads -- spanning transformer-based models and attention-free CNNs -- and across three floating-point formats. Our central empirical finding is a sharp bit-sensitivity transition: flipping any of the least-significant fraction bits up to a data-type-specific threshold, Xsafe, degrades task metrics by less than 1% under deterministic single-bit stress tests. Sensitivity rises through the upper fraction bits and spikes at the exponent-mantissa boundary, where a single-bit flip causes catastrophic collapse. Because low-order bits are largely inconsequential while high-order and exponent bits are critical, uniform SECDED protection -- which guards every bit equally at 12.5% storage overhead -- is unnecessarily conservative. We derive per-data-type Xsafe floors (FP16: 6, BF16: 4, FP32: 15) and workload-aware tiers that widen the unprotected region for resilient model classes, raising ECC savings to 37.5-62.5% without retraining. Text-conditioned diffusion models dictate the conservative floor; vision encoders, NLU models, and resilient LLMs tolerate wider bypass regions. These floors and tiers drive an Unequal Error Protection (UEP) codec with per-cacheline data-type tags and a dual-partition SRAM architecture for ML accelerators. Validation across 870+ fault-injection runs confirms selective protection holds under contiguous 2- and 3-bit upsets. The codec reduces ECC area by 27.8% relative to uniform SECDED; dual-voltage operation of the non-critical partition lowers gross BF16 read energy by about 17%, with a roughly 4% dual-partition macro-area overhead.
CommentsAccepted to appear at the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026). 14 pages