arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度微压缩:面向微控制器的结构化剪枝与位打包量化

Deep Microcompression: Structured Pruning and Bit-packed Quantization for Microcontrollers

Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe

arXiv 2609.05081首次发表:更新:

发表机构

PowerLabs Technologies(PowerLabs科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出面向裸机微控制器的硬件感知深度学习压缩流水线Deep Microcompression,整合结构化剪枝等技术,在LeNet-5上实现高压缩比,还实现了标准CNN在2KB SRAM的ATmega328P上的部署。

AI 中文摘要

本文提出Deep Microcompression(DMC),一种针对裸机微控制器深度学习推理的硬件感知流水线。DMC整合结构化剪枝、量化感知训练与定长位打包,在LeNet-5(准确率98.77%)上实现55.8倍权重压缩比,生成无依赖、延迟确定的C库。在RP2040(Cortex-M0+)上,DMC相比TensorFlow Lite将二进制大小减少3倍且匹配其准确率,关键是DMC实现了首个在仅2KB SRAM的ATmega328P上部署标准CNN的记录,此前该设备被认为无法进行CNN推理。

英文摘要

This paper introduces Deep Microcompression (DMC), a hardware-aware pipeline for deep learning inference on bare-metal microcontrollers. DMC integrates structured pruning, quantization-aware training, and fixed-length bit-packing to achieve a 55.8$\times$ weight compression ratio on LeNet-5 (98.77\% accuracy), generating a dependency-free C library with deterministic latency. On the RP2040 (Cortex-M0+), DMC reduces binary size by 3$\times$ versus TensorFlow Lite while matching its accuracy. Critically, DMC enables the first documented deployment of a standard CNN on the ATmega328P, a device constrained to 2KB SRAM, previously considered infeasible for CNN inference.

CommentsPresented at the Global South ML Workshop at the International Conference on Machine Learning (ICML 2026), Seoul, South Korea

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑