发表机构
PowerLabs Technologies(PowerLabs科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出面向裸机微控制器的硬件感知深度学习压缩流水线Deep Microcompression,整合结构化剪枝等技术,在LeNet-5上实现高压缩比,还实现了标准CNN在2KB SRAM的ATmega328P上的部署。
AI 中文摘要
本文提出Deep Microcompression(DMC),一种针对裸机微控制器深度学习推理的硬件感知流水线。DMC整合结构化剪枝、量化感知训练与定长位打包,在LeNet-5(准确率98.77%)上实现55.8倍权重压缩比,生成无依赖、延迟确定的C库。在RP2040(Cortex-M0+)上,DMC相比TensorFlow Lite将二进制大小减少3倍且匹配其准确率,关键是DMC实现了首个在仅2KB SRAM的ATmega328P上部署标准CNN的记录,此前该设备被认为无法进行CNN推理。
英文摘要
This paper introduces Deep Microcompression (DMC), a hardware-aware pipeline for deep learning inference on bare-metal microcontrollers. DMC integrates structured pruning, quantization-aware training, and fixed-length bit-packing to achieve a 55.8$\times$ weight compression ratio on LeNet-5 (98.77\% accuracy), generating a dependency-free C library with deterministic latency. On the RP2040 (Cortex-M0+), DMC reduces binary size by 3$\times$ versus TensorFlow Lite while matching its accuracy. Critically, DMC enables the first documented deployment of a standard CNN on the ATmega328P, a device constrained to 2KB SRAM, previously considered infeasible for CNN inference.
CommentsPresented at the Global South ML Workshop at the International Conference on Machine Learning (ICML 2026), Seoul, South Korea