arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

渐进式多祖先位深蒸馏

Progressive Multi-Ancestor Bit-Depth Distillation

Adil Mubashir Chaudhry, Osama Ahmad, Zubair Khalid, Murtaza Taj

arXiv 2610.04100首次发表:更新:

发表机构

LUMS School of Science and Engineering; University of Massachusetts Amherst(拉合尔管理科学大学科学与工程学院; 马萨诸塞大学阿默斯特分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出渐进式多祖先位深蒸馏(PMABD)框架,通过多祖先多阶段蒸馏稳定超低位量化,在多个数据集上优于现有压缩方法,性能提升1.06%。

AI 中文摘要

模型压缩策略被广泛用于减少内存占用和网络复杂度,特别是对于计算、内存和能源资源受限的设备。先前依赖从浮点高精度(FP32)到整数低精度(INT4)表示的同步转换并蒸馏到较小模型的工作,存在训练不稳定和预测性能急剧下降的问题。为解决这些限制,我们提出了一个统一框架,称为渐进式多祖先位深蒸馏(PMABD),该框架在通过不断增长的高精度祖先教师池传递知识的同时,逐步压缩网络。PMABD生成一系列中间教师,每个教师都从所有更高精度的祖先学习,并共同监督最终的目标学生。这种多祖先、多阶段的设计通过降低训练过程中的量化噪声分布并确保稳定的量化,稳定了超低位量化。在CIFAR-10/100上使用ResNet-20/32/18,以及在Tiny-ImageNet上使用MobileNetV2的实验表明,PMABD优于最先进的压缩框架,使W2A2(ResNet-18/CIFAR-100)学生模型的性能提高了1.06%。我们表明,基于饱和的停止标准有助于提高最终学生的性能。

英文摘要

Model compression strategies are widely employed to reduce memory footprint and network complexity, particularly for devices with constrained computational, memory, and energy resources. Prior works that rely on simultaneous conversion from floating-point high-precision (FP32) to integer low-precision (INT4) representations and distillation into smaller models suffer from unstable training and drastic degradation of prediction performance. To address these limitations, we propose a unified framework, known as \textbf{P}rogressive \textbf{M}ulti-\textbf{A}ncestor \textbf{B}it-depth \textbf{D}istillation (PMABD), that progressively compresses the network while transferring knowledge through a growing pool of higher-precision ancestor teachers. PMABD generates a sequence of intermediate teachers that each learn from all higher-precision ancestors and jointly supervise the final target student. This multi-ancestor, multi-stage design stabilizes ultra-low-bit quantization by lowering quantization noise profiles across training and ensuring stable quantization. Experiments on CIFAR-10/100 with ResNet-20/32/18, and Tiny-ImageNet with MobileNetV2 show that PMABD outperforms state-of-the-art compression frameworks, results in 1.06$\%$ increase in performance of W2A2 (ResNet-18/CIFAR-100) student model. We show that a saturation-based stopping criterion contributes to improve the performance of our final student.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑