发表机构
Johns Hopkins University; KU Leuven(约翰斯·霍普金斯大学; 鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ShatterQuant通过软硬件协同设计,在脉动阵列加速器上实现块级混合精度量化,在低2位有效位宽下达到接近最先进精度,并提升面积与能效。
AI 中文摘要
由于传统加速器对张量内部异构精度的支持有限,神经网络量化在很大程度上仍局限于逐张量的精度分配。我们提出ShatterQuant,一个软硬件协同设计的框架,通过为权重投影的各个块分配独立的位宽,实现在每个张量内部的混合精度量化。ShatterQuant将精度粒度与PE配置相耦合,使得每种精度决定一个有效的块高度。我们引入:(1)一种硬件感知的训练后方法,基于块级标准差和权重敏感性分配张量内部精度;(2)ShatterQuant Transformer加速器,支持1/2/4/8位权重精度、依赖精度的PE配置、块重缩放,以及集成的softmax和分段线性非线性函数;(3)使用TSMC 16nm PDK在1 GHz频率下运行的实现来评估模型-硬件权衡,达到1.5 TOPS、760 GOPS/$mm^2$面积效率和2.8 TOPS/W能效。在DeiT和ImageNet-1K上,ShatterQuant在有效位宽低2位的情况下,精度达到最先进混合精度技术的$3.3\%$以内,而在PixelDiT上展示了相当的生成质量。ShatterQuant展示了如何通过软硬件协同设计实现细粒度的张量内部混合精度。
英文摘要
Due to limited support for intra-tensor heterogeneous precision in conventional accelerators, neural network quantization remains largely restricted to per-tensor precision assignment. We present ShatterQuant, a hardware-software co-designed framework enabling mixed-precision quantization within each tensor by assigning independent bit-widths to blocks of a weight projection. ShatterQuant couples precision granularity with PE configuration, such that each precision determines an effective block height. We introduce (1) a hardware-aware post-training method that assigns intra-tensor precision based on block-level standard deviation and weight sensitivity; (2) the ShatterQuant Transformer Accelerator supporting 1/2/4/8-bit weight precision, precision-dependent PE configuration, block rescaling, and integrated softmax and piecewise-linear nonlinearities; and (3) an evaluation of model-hardware tradeoffs using an implementation in the TSMC 16nm PDK operating at 1 GHz, achieving 1.5 TOPS, 760 GOPS/$mm^2$ area efficiency, and 2.8 TOPS/W energy efficiency. On DeiT and ImageNet-1K, ShatterQuant achieves accuracy within $3.3\%$ of state-of-the-art mixed-precision techniques while using a 2 bit lower effective bitwidth, while for PixelDiT demonstrates comparable generation quality. ShatterQuant demonstrates how fine-grained intra-tensor mixed-precision can be realized through hardware-software co-design.