arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

二值化神经网络的剪枝:专用框架与全局加权算法

Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms

Roan Rubiales, Jean Pierre David

arXiv 2608.26233首次发表:更新:

发表机构

Polytechnique Montreal(蒙特利尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对二值化神经网络剪枝适配问题,提出基于PyTorch的专用框架及全局加权剪枝算法,在VGG11二值化模型上实现70%剪枝率且准确率不变,优于现有最优结果。

AI 中文摘要

深度神经网络的极致压缩,直至完全二值化,可大幅减少内存占用与算术复杂度,便于在带有现场可编程门阵列(FPGAs)和微控制器的受限边缘硬件上部署。尽管将二值化与剪枝结合有望进一步提升效率,但现有剪枝策略并不适配二值化表示,且极少能转化为实际硬件节省。我们提出一种基于PyTorch、面向研究的框架,该框架集成了冻结与剪枝机制,用于设计和优化二值化神经网络。此框架支持对现有最优方法进行快速可复现评估,并能快速原型化新方法。借助该框架,我们提出一种新型剪枝方法,该方法考虑了不同抽象层级上学到参数的相对重要性。这种全局加权机制在模型准确率与剪枝率之间始终实现更优权衡,在VGG11上达到70%的剪枝率且准确率保持不变,而在二值化设置下,现有最优结果仅达到41%。

英文摘要

Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontrollers. Although combining binarization with pruning promises additional efficiency gains, existing pruning strategies are ill-suited to binarized representations and rarely translate into meaningful hardware savings. We introduce a PyTorch-based, research-oriented framework that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networks. The framework enables rapid and reproducible evaluation of state-of-the-art approaches and the fast prototyping of new ones. Leveraging this framework, we propose a novel pruning method that accounts for the relative importance of learned parameters across abstraction levels. Such a global weighting mechanism consistently achieves a superior trade-off between model accuracy and pruning rate, achieving a 70% pruning rate on VGG11 with constant accuracy, while state-of-the-art results reach only 41% in the binarized setting.

Comments9 pages, 3 figures, 5 tables, 3 algorithms

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑