arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NANQ:面向模拟内存计算的噪声底感知混合精度非均匀量化

NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory

Yizhe Chen, Wenshuai Yao, Saiya Wang, Yuannuo Feng, Wenbo Qi, Kechao Tang, Ngai Wong, Wenyong Zhou, Wang Kang

arXiv 2608.02700首次发表:更新:

AI 中文总结

针对模拟内存计算的噪声问题,提出NANQ量化框架,实现混合精度非均匀量化,在2比特量化下提升视觉模型准确率、降低语言模型PPL,仅用少量等效比特即可获显著增益。

AI 中文摘要

模拟内存计算(CIM)可实现高能效的神经网络推理,但器件变异与读取噪声会严重降低低比特量化模型的性能。现有面向CIM的量化方法主要最小化理想量化误差,忽略硬件噪声底,导致精度分配低效。本文提出NANQ,一种面向模拟CIM的噪声感知混合精度非均匀量化框架。NANQ基于eFlash CIM阵列实测响应建模与幅度相关的权重噪声,将噪声分布转换为自适应量化密度,对低噪声区域分配更精细的分辨率,同时避免在噪声主导区域采用无效精度;还通过统一阈值识别各层在硬件噪声下的精度饱和点,分配层级比特宽。在eFlash CIM SoC上的片上实验表明,在2比特权重幅度量化下,NANQ相较于PowerQuant,视觉模型准确率提升8.05个百分点,语言模型困惑度(PPL)平均降低54.7%;混合精度NANQ仅用3.2-3.8等效比特,即可捕获额外量化资源可实现的大部分增益。

英文摘要

Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit quantized models. Existing CIM-oriented quantization methods mainly minimize ideal quantization error, ignoring the hardware noise floor and thus causing inefficient precision allocation. We propose NANQ, a noise-aware mixed-precision non-uniform quantization framework for analog CIM. NANQ models magnitude-dependent weight noise from measured responses of an eFlash CIM array and converts the noise profile into an adaptive quantization density, assigning finer resolution to low-noise regions while avoiding ineffective precision in noise-dominated regions. It further assigns layer-wise bit-widths by identifying each layer's precision saturation point under hardware noise using a unified threshold. On-chip experiments on an eFlash CIM SoC show that, under 2-bit weight-magnitude quantization, NANQ improves vision-model accuracy by 8.05 percentage points and reduces language-model PPL by 54.7% on average over PowerQuant. Mixed-precision NANQ captures most of the gains obtainable from additional quantization resources with only 3.2-3.8 equivalent bits.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑