发表机构
Intel(英特尔)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出循环残差量化(RRQ)框架,可从单个检查点生成LLM的多种精度表示,构建速度比MatGPTQ快3.3倍,在6、8比特精度下表现有竞争力,代码将公开。
AI 中文摘要
在各类部署约束下服务大语言模型(LLM),需在精度、内存占用和吞吐量间灵活权衡,但传统量化方法通常需为每个目标比特宽度单独准备检查点。本文提出循环残差量化(Recurrent Residual Quantization, RRQ),这是一种后训练量化(PTQ)框架,将权重表示为低比特量化基与一系列量化残差修正项,可从单个检查点生成多种有效精度。RRQ从通过后训练量化(PTQ)或舍入取最近(RTN)得到的2比特模型出发,逐步添加通过RTN生成的轻量2比特残差,构建4、6、8比特表示。该方法无需校准,避免了联合多比特优化。在Qwen3-8B设置下,完整的全RTN 2/4/6/8比特包构建耗时1293秒,比实测的MatGPTQ构建速度快3.3倍。对6款近期大语言模型的实验显示,其在6、8比特精度下表现有竞争力,在4比特下则呈现模型依赖的特性,代码将在发表后公开。
英文摘要
Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional quantization methods typically require a separate checkpoint for each target bit-width. We introduce Recurrent Residual Quantization (RRQ), a post-training quantization (PTQ) framework that represents weights as a low-bit quantized base together with a sequence of quantized residual corrections, enabling multiple effective precisions from a single checkpoint. Starting from a 2-bit model obtained via post-training quantization (PTQ) or round-to-nearest (RTN), RRQ progressively adds lightweight 2-bit residuals generated via RTN to construct 4-, 6-, and 8-bit representations. The method is calibration-free and avoids joint multi-bit optimization. In our Qwen3-8B setup, the full all-RTN 2-/4-/6-/8-bit package is constructed in 1,293 seconds, 3.3 times faster than the measured MatGPTQ construction. Experiments on six recent LLMs show competitive accuracy at 6 and 8 bits, with model-dependent behavior at 4 bits. The code will be made publicly available upon publication.
CommentsNeurIPS 2026 submission; 14 tables, 1 algorithm, and no figures