发表机构
Bielefeld University(比勒费尔德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出W16A16,一种16位高精度量化方法,在Armv7E-M微控制器上以8位成本实现约10倍更低的量化误差,同时保持相似或更优的速度与能耗。
AI 中文摘要
为了在边缘硬件上部署深度神经网络,需要高效且保持高准确率的推理方案。本文提出了W16A16,一种高精度(16位)、高速、低能耗的量化方法。在广泛应用的微控制器架构Armv7E-M上,与替代量化方案相比,我们提出的方法在层级别和模型级别均实现了更快的速度和更低的能耗。我们分析了Armv7E-M的架构,解释了16位方法性能优势背后的基本原理,并评估了回归和分类任务的实证量化误差,以及MCU部署中的实证时间和能耗。我们观察到,与8位量化方案相比,量化误差降低了约10倍,同时实现了相似或更好的推理时间和能耗。
英文摘要
To deploy deep neural networks on edge hardware, highly efficient inference schemes are necessary that retain high accuracy. This work presents W16A16, a high precision (16-bit), fast speed, low energy quantization method. On a widely applied microcontroller architecture Armv7E-M, our proposed approach achieves faster speed and lower energy consumption on layer- and model-level compared to alternative quantization schemes. We analyze the architecture of Armv7E-M, explain the underlying principles behind the performance advantages of 16-bit approaches, and evaluate the empiric quantization errors for regression and classification tasks, as well as empiric time- and energy consumption in MCU deployment. We observe ca.\ 10 times lower quantization errors compared to 8-bit quantization schemes while achieving similar or better inference times and energy consumption.