发表机构
National University of Singapore; Central South University(新加坡国立大学; 中南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于码本的LLM量化中码字分配不均衡问题,提出BARQ方法,通过熵正则化最优传输实现均衡软分配并精炼码本,在多个模型上以更低困惑度和更高零样本准确率超越基线。
AI 中文摘要
随着大语言模型(LLM)参数数量的增长,模型存储和参数内存流量已成为高效部署的主要瓶颈。基于码本的权重量化降低了这些成本,但在拟合过程中,最近码字分配的不均衡可能导致某些码字更新不足,限制了码本的有效利用率。为解决这一局限,我们提出了均衡分配精炼量化(BARQ),通过均衡拟合现有码本来提高量化质量。具体而言,我们首先通过熵正则化的最优传输(具有均匀边际和曲率加权重建成本)计算权重块与码字之间的联合软分配,确保在精确解中每个码字获得相等的正拟合质量。然后,我们通过分配加权的重心更新来精炼码字,并证明该更新在固定分配下最小化拟合目标。对于有限的Sinkhorn迭代,只要所有码字质量超过分母下限,所实现的更新保持此最优性。最后,我们丢弃软分配,使用精炼后的码本进行标准的硬最近码字编码,我们的分析建立了减少硬量化失真和评估损失的充分条件。在多个LLM上,BARQ在相当的比特预算下,相比评估的基线实现了更低的困惑度和更高的平均零样本准确率。代码可在该https URL获取。
英文摘要
As large language models (LLMs) grow in parameter count, model storage and parameter memory traffic have become major bottlenecks to efficient deployment. Codebook-based weight quantization reduces these costs, but imbalanced nearest-codeword assignments during fitting can leave some codewords insufficiently updated, limiting effective codebook utilization. To address this limitation, we propose Balanced Assignment Refinement for Quantization (BARQ), which improves quantization quality through balanced fitting of existing codebooks. Specifically, we first compute joint soft assignments between weight blocks and codewords through entropically regularized optimal transport with uniform marginals and curvature-weighted reconstruction costs, ensuring equal positive fitting mass for every codeword in the exact solution. We then refine the codewords through an assignment-weighted barycentric update, which we prove minimizes the fitting objective for fixed assignments. For finite Sinkhorn iterations, the implemented update retains this optimality provided all codeword masses exceed the denominator floor. Finally, we discard the soft assignments and use the refined codebook for standard hard nearest-codeword encoding, with our analysis establishing sufficient conditions for reducing hard-quantization distortion and evaluation loss. Across multiple LLMs, BARQ achieves lower perplexity and higher mean zero-shot accuracy than the evaluated baselines at comparable bit budgets. The code is available at https://github.com/chenhangcuisg-code/BARQ.
CommentsCode: https://github.com/chenhangcuisg-code/BARQ