发表机构
School of Computer Science, Peking University; School of Software and Microelectronics, Peking University; School of Physics, Peking University; The Chinese University of Hong Kong, Shenzhen; Central Research Institute, ZTE Corporation(北京大学计算机学院; 北京大学软件与微电子学院; 北京大学物理学院; 香港中文大学(深圳); 中兴公司中央研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReQuant是用于训练后量化(PTQ)的无反向传播固定网格细化方法,可作为即插即用后处理阶段,提升各类PTQ初始化器生成的量化模型性能,在低位宽下增益显著。
AI 中文摘要
训练后量化(PTQ)被广泛用于降低大语言模型的内存和计算成本。现有PTQ方法通常通过启发式规则或贪婪优化获得初始量化模型,量化完成后得到的整数赋值通常被视为最终结果。这一观察结果促使我们在PTQ中引入一个补充优化阶段:在生成可执行的量化模型后,保持量化格式不变,同时让量化权重仍可优化。我们提出ReQuant,这是一个用于该阶段的无反向传播固定网格细化过程。ReQuant不依赖PTQ初始化器,以现有量化模型为可行起点,在固定量化网格上迭代修正其离散权重赋值。被接受的更新会严格降低均方重构误差,且始终保持在原始网格上。通过这种方式,ReQuant将初始固定的PTQ输出转化为可迭代优化的离散解,可作为现有PTQ流程的即插即用后处理阶段。在不同模型家族、位宽和下游任务上的实验表明,ReQuant能持续改进来自异构PTQ初始化器的量化模型,在简单初始化器和较低位宽下增益尤为显著。值得注意的是,ReQuant可对简单的四舍五入到最近值的初始化进行多轮细化,直至在相同量化格式下接近或超越GPTAQ。这些结果确立了ReQuant作为进一步改进现有PTQ流程的实用补充阶段的地位。
英文摘要
Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically obtain an initial quantized model through heuristic rules or greedy optimization, and once quantization is completed the resulting integer assignments are usually treated as final. This observation motivates a complementary optimization stage within PTQ that keeps quantized weights improvable after an executable quantized model has been produced, while preserving the quantized format. We introduce ReQuant, a backpropagation-free fixed-grid refinement procedure for this stage. Agnostic to the PTQ initializer, ReQuant takes an existing quantized model as a feasible starting point and iteratively revisits its discrete weight assignments on the fixed quantization grid. Accepted updates strictly reduce the mean squared reconstruction error and remain on the original grid. In this way, ReQuant turns the initially fixed PTQ output into an iteratively optimizable discrete solution and serves as a plug-and-play post-processing stage for existing PTQ pipelines. Experiments across diverse model families, bit-widths, and downstream tasks show that ReQuant consistently improves quantized models from heterogeneous PTQ initializers, with especially large gains on simple initializers and lower bit-widths. Notably, ReQuant can refine a simple round-to-nearest initialization across multiple sweeps until it approaches or surpasses GPTAQ under the same quantization format. These results establish ReQuant as a practical complementary stage for further improving existing PTQ pipelines.
Comments10 pages, 3 figures, 16 tables