发表机构
Univ Rennes; Inria; CNRS; IRISA(雷恩大学; 法国国家信息与自动化研究所; 法国国家科学研究中心; 法国雷恩信息系统与随机系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MicroQonv通过单次量化和改进的im2col,高效实现卷积层微缩放,降低量化成本与内存移动,提升边缘持续学习准确性。
AI 中文摘要
微缩放量化技术越来越多地用于以8位或更少的位数表示神经网络参数,同时保持接近全精度的准确性。然而,在卷积层中高效应用这些方法并不简单。一种朴素的方法是将全精度权重和激活值传输到处理单元,并对每个张量进行两次量化,导致比预期多得多的内存移动。额外的开销来自激活张量,由于在量化前应用了im2col变换,其尺寸大幅增长。我们提出了MicroQonv,一种将微缩放与卷积层的前向和反向操作相结合的方法,通过仅对每个张量量化一次,并在应用修改版的im2col(即通道-批次优先im2col)之前对激活张量进行量化。MicroQonv将权重和梯度的量化成本降低了×2倍,激活的量化成本降低了高达×9倍,且准确性损失可忽略不计。与全精度对应物相比,它减少了高达×7.53倍的内存移动和存储。这样,MicroQonv将最先进的目标检测模型YOLOV8nano的微缩放量化激活内存移动减少了×3.5倍,YOLOV26nano减少了×2.2倍。它还在边缘持续学习的量化潜在重放策略中实现了4位微缩放,将准确性提高了+5.7%至+11%。
英文摘要
Microscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full precision accuracy. However, applying these methods efficiently in convolutional layers is not straightforward. A naive approach transfers full-precision weights and activations to processing units and quantizes each tensor twice, resulting in much more memory movement than expected. Additional overhead comes from the activation tensors, whose sizes grow substantially because of the im2col transformation applied before quantization. We propose MicroQonv, a way to combine microscaling with convolutional layers' forward and backward operations by quantizing each tensor only once and quantizing the activation tensor before applying a modified version of im2col: channel-batch-first im2col. MicroQonv reduces the quantization cost by a factor of $\times2$ for weights and gradients, and by up to $\times9$ for activations, at a negligible accuracy cost. It reduces memory movement and storage by up to $\times7.53$ compared to their full-precision counterparts. This way, MicroQonv reduces microscaling-quantized activation memory movement by $\times3.5$ for state-of-the-art object detection models YOLOV8nano and $\times2.2$ for YOLOV26nano. It also enables 4-bit microscaling in a quantized latent replay strategy for continual learning at the edge, improving accuracy by +5.7% to +11%.
Comments12 pages, 7 figures