arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12111cs.ET

RAPID-DNN:面向边缘AI的鲁棒DNN推理部署的可靠性感知划分

RAPID-DNN: Reliability-Aware Partitioning for Robust DNN Inference Deployment on Edge AI

Mukta Debnath, Krishnendu Guha, Debasri Saha, Amlan Chakrabarti, Susmita Sur-Kolay

首次发表
浏览论文内容

中文总结 AI 辅助

RAPID-DNN是将可靠性作为第三优化目标的DNN划分框架,通过NSGA-II优化器在AlexNet等模型上,相比基线提升了Top-1准确率并降低了故障诱导的准确率下降,实现了边缘加速器上的鲁棒DNN部署。

中文摘要 AI 辅助

部署在异构边缘加速器上的深度神经网络(DNN)日益面临运行时硬件故障的威胁。严格的功耗和资源预算迫使这些平台采用激进的电压缩放,并限制硬件故障缓解措施的使用,而工艺偏差、热应力和器件老化进一步提高了故障发生率。由此产生的瞬态比特级损坏会在推理过程中传播,可能大幅降低预测准确率。现有的DNN划分框架在无故障假设下优化延迟和能耗,因此其部署策略在实际运行条件下往往无法维持可靠的推理。本文提出RAPID-DNN,这是一个可靠性感知划分框架,将运行时可靠性作为与延迟、能耗并列的第三个优化目标。RAPID-DNN首先通过在权重和激活域中进行系统的故障注入来表征层级可靠性,以识别脆弱性关键层。所得的敏感性曲线指导三目标NSGA-II优化器,该优化器共同最小化推理延迟、能耗以及预期故障诱导的准确率下降。我们使用异构Eyeriss和SIMBA加速器配置文件,在AlexNet、SqueezeNet、ResNet18、VGG16和MobileNetV2上评估RAPID-DNN。在不同故障模型和注入概率下,它始终比不考虑故障的基线提高推理鲁棒性。在代表性的混合随机故障配置下,RAPID-DNN将平均Top-1准确率提高9.81个百分点,将平均预期故障诱导的准确率下降降低36.86%,代价是平均2.38%的延迟开销和3.34%的能耗开销。这些结果表明,将可靠性作为划分目标可实现异构边缘加速器在运行时硬件故障下的鲁棒DNN部署。

英文摘要

Deep neural networks (DNNs) deployed on heterogeneous edge accelerators are increasingly exposed to runtime hardware faults. Tight power and resource budgets push these platforms toward aggressive voltage scaling and limit the use of hardware fault mitigation, while process variation, thermal stress, and device aging raise fault rates further. The resulting transient bit level corruptions propagate through inference and can sharply reduce prediction accuracy. Existing DNN partitioning frameworks optimize latency and energy under fault free assumptions, so their deployment strategies often fail to sustain reliable inference in realistic operating conditions. This paper presents RAPID DNN, a reliability aware partitioning framework that adds runtime reliability as a third optimization objective alongside latency and energy. RAPID DNN first characterizes layer wise reliability through systematic fault injection in the weight and activation domains to identify vulnerability critical layers. The resulting sensitivity profile guides a three objective NSGA II optimizer that jointly minimizes inference latency, energy consumption, and expected fault induced accuracy degradation. We evaluate RAPID DNN on AlexNet, SqueezeNet, ResNet18, VGG16, and MobileNetV2 using heterogeneous Eyeriss and SIMBA accelerator profiles. Across fault models and injection probabilities, it consistently improves inference robustness over the fault agnostic baseline. Under the representative mixed random fault configuration, RAPID DNN raises average Top 1 accuracy by 9.81 percentage points and reduces average expected fault induced accuracy degradation by 36.86%, at a cost of 2.38% average latency overhead and 3.34% average energy overhead. These results show that treating reliability as a partitioning objective enables robust DNN deployment on heterogeneous edge accelerators under runtime hardware faults.

发表机构

  • A.K. Choudhury School of Information Technology, University of Calcutta(加尔各答大学 A.K. Choudhury 信息技术学院)
  • School of Computer Science and Information Technology, University College Cork(科克大学学院 计算机科学与信息技术学院)
  • Indian Institute of Technology, Kharagpur(印度理工学院 卡拉格普尔分校)
  • Indian Statistical Institute(印度统计研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑