arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06781cs.AR

BLINK:基于批归一化的完整性检查点,用于深度神经网络加速器中多种权重损坏的原位检测与缓解

BLINK: Batch Normalization-based Integrity Checkpoints for In-Situ Detection and Mitigation of Diverse Weight Corruptions in DNN Accelerators

  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

Marzia Khan, Akul Malhotra, Sumeet Kumar Gupta

AI总结:

BLINK利用批归一化统计量偏移检测并缓解DNN加速器中的多种权重损坏,实现高精度检测与精度恢复,且硬件开销极低。

AI中文摘要:

在安全关键型部署中,AI硬件必须能够抵御各种威胁,如老化、软错误、硬故障和对抗性攻击(例如渐进式位翻转攻击(PBFA))。所有这些威胁都会破坏存储的权重,而芯片仍会继续产生自信但不准确的预测。检测和缓解此类权重扰动对于安全关键型平台至关重要。为此,我们提出了BLINK,一种基于片上批归一化(BN)的即时检测与缓解方法,它基于对激活统计量偏移的持续感知,针对多种权重损坏(随机和局部故障以及对抗性位翻转)进行检测。BLINK分两个阶段运行:(1)离线预表征激活偏移与推理精度下降之间的关系;(2)片上运行时检测和缓解权重损坏。检测到损坏后,被标记的层会在同一次前向传播中重新居中,使其更接近存储的干净参考。BLINK完全自主,无需主机通信、操作暂停或访问微调数据。如果缓解后的残余偏移表明精度已降至用户设定的下限以下,则一个保留的监视器会中止推理。在ResNet-20/50和MobileNetV2上针对CIFAR-10/100进行评估,BLINK对所有故障类型的检测精度均超过99%。此外,在0.5%随机位翻转(ResNet-50/CIFAR-10)下,它可将精度从10%恢复到85.88%;对于局部故障(MobileNetV2/CIFAR-10),可恢复到84%;在PBFA(ResNet-20/CIFAR-10)下,可从随机猜测精度恢复到80%-83%。硬件开销估算表明,BLINK产生的成本可忽略不计,延迟增加不到2%,计算开销仅增加0.53%。

英文摘要:

In safety-critical deployments, AI hardware must remain reliable against a broad spectrum of threats such as aging, soft errors, hard faults, and adversarial attacks (e.g. progressive bit flip attack (PBFA)). All of these corrupt stored weights while the chip keeps producing confident but inaccurate predictions. Detecting and mitigating such weight perturbations is crucial for safety-critical platforms. To that end, we propose BLINK, an on-chip batch normalization (BN)-based on-the-fly detection and mitigation approach, which is based on continual sensing of the shift in the activation statistics, and targets a wide variety of weight corruptions (random and localized faults as well as adversarial bit flips). BLINK operates in two phases: (1) off-line pre-characterization of the relationship of the activation shifts with inference accuracy drop, and (2) on-chip runtime detection and mitigation of weight corruptions. Upon detection, the flagged layer is re-centered to bring it closer to its stored clean reference within the same forward pass. BLINK is fully autonomous, eliminating the need for host communication, operation halts, or access to fine-tuning data. If the residual shift after mitigation indicates that accuracy has fallen below a user-set floor, a held-out watcher aborts the inference. Evaluated on ResNet-20/50 and MobileNetV2 for CIFAR-10/100, BLINK detects harmful corruptions with >99% precision across all fault types. Further, it recovers accuracy from 10% to 85.88% under 0.5% random bit flips (Resnet-50/CIFAR-10), up to 84% for localized faults (MobileNetV2/CIFAR-10), and from random-guess accuracy to 80%-83% under PBFA (ResNet-20/CIFAR-10). Hardware overhead estimates indicate that BLINK incurs negligible costs, with less than a 2% increase in latency and only a 0.53% increase in computation overhead.

↑