REQAP:面向边缘DNN加速的弹性权重打包与量化
REQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration
- Humboldt University of Berlin(柏林洪堡大学)
- Tallinn University of Technology(塔林理工大学)
- University of Zanjan(赞詹大学)
- Brandenburg University of Technology Cottbus-Senftenberg(勃兰登堡工业大学科特布斯-森夫滕贝格分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出REQAP方法,通过敏感度驱动的混合精度量化与寄存器级打包,结合MSB复制实现TMR保护,在边缘DNN加速器上显著降低内存和计算开销,并提升故障下的精度弹性。
AI中文摘要:
在边缘加速器上高效部署深度神经网络(DNN)需要在保持可靠性的同时进行积极的模型压缩,尤其是在易发生故障的硬件环境中。本文提出了一种面向基于脉动阵列的DNN加速器的可靠性感知量化权重打包方法。一个基于敏感度的混合精度量化框架根据精度影响分配逐层位宽,同时强制权重与激活之间的对称精度。一种确定性的寄存器级打包策略将多个异构操作数对整合到固定宽度的寄存器字中,实现寄存器内SIMD(SWAR)风格的并行执行,从而减少内存占用和执行周期。为了提高对硬件故障的弹性,选择性位级保护将关键层的最高有效位(MSB)复制到未使用的寄存器空间中,以最小开销实现三模冗余(TMR)风格的保护。开发了一个脉动阵列仿真框架,以在真实执行条件下评估所提出的打包和容错机制。在AlexNet、VGG-11和ResNet-18上的仿真表明,内存最多减少62%,乘累加(MAC)操作最多减少56%,同时在故障注入下相比基线和完全保护的模型显著提高了精度弹性。
英文摘要:
Efficient deployment of Deep Neural Networks (DNNs) on edge accelerators requires aggressive model compression while maintaining reliability in fault-prone hardware environments. This paper presents a reliability-aware quantized weight packing methodology for systolic-array-based DNN accelerators. A sensitivity-driven mixed-precision quantization framework assigns layer-wise bit-widths according to accuracy impact while enforcing symmetric precision between weights and activations. A deterministic register-level packing strategy consolidates multiple heterogeneous operand pairs into fixed-width register words, enabling SIMD-within-a-register (SWAR) style parallel execution that reduces both memory footprint and execution cycles. To improve resilience against hardware faults, selective bit-level protection replicates the most significant bits (MSBs) of critical layers into unused register space, achieving TMR-style protection with minimal overhead. A systolic-array simulation framework is developed to evaluate the proposed packing and fault-tolerance mechanisms under realistic execution conditions. Simulations in AlexNet, VGG-11, and ResNet-18 demonstrate up to 62% memory reduction and up to 56% reduction in Multiply-Accumulate (MAC) operations, while significantly improving accuracy resilience under fault injection compared to baseline and fully protected models.