arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少比特,一法则:迈向 W2A4KV2

Few Bits, One Law: Toward W2A4KV2

Kai Yi, Tarek Elgamal, Sruthikesh Surineni, Vignesh Vivekraja, Soumyadeep Ghosh, Steven Li

arXiv 2610.09202首次发表:更新:

发表机构

Meta AI(Meta AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CanonQ统一量化感知训练框架,通过源规范化和任务适应分离,在W2A4KV2联合压缩下显著降低困惑度并提升零样本准确率,同时惠及代码生成与数学推理任务。

AI 中文摘要

当权重、激活和KV缓存被同时量化时,极端低比特LLM压缩最具挑战性:它们的分布不同,且量化误差在整个网络中相互作用。我们提出CanonQ,一个统一的量化感知训练框架,通过将源规范化与任务感知适应分离来解决这些挑战。固定的旋转和能量归一化将异构张量源映射到规范坐标,使得冻结的高斯参考码本能够在不同层和模型间复用。随后,联合训练在统一的标量/向量接口内,使网络适应权重、激活和缓存量化的耦合误差。我们界定了冻结码本迁移误差和局部任务损失,并推导出一个精确的、归一化感知的直通雅可比矩阵,将量化失真与梯度偏差联系起来。最显著的收益出现在联合W2A4KV2压缩下:在LLaMA3-1B/3B/8B上,CanonQ-Omni相比先前最先进方法和代表性量化基线,实现了WikiText-2困惑度最高降低14.28倍,平均零样本准确率最高提升57.9%。这些优势扩展到Qwen3-1.7B、代码生成和数学推理:在指令微调的MobileLLM-Pro-1B上,W2A16KV16设置下,CanonQ在HumanEval pass@1上相对最强评估量化基线取得41.7%的相对提升,在GSM8K精确匹配上取得39.1%的相对提升。

英文摘要

Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local task loss, and derive an exact normalization-aware straight-through Jacobian that links quantization distortion to gradient bias. The strongest gains arise under joint W2A4KV2 compression: across LLaMA3-1B/3B/8B, CanonQ-Omni achieves up to 14.28x lower WikiText-2 perplexity and up to 57.9% higher mean zero-shot accuracy than prior state-of-the-art and representative quantization baselines. The benefits extend to Qwen3-1.7B, code generation, and mathematical reasoning: on instruction-tuned MobileLLM-Pro-1B at W2A16KV16, CanonQ achieves relative improvements of 41.7% in HumanEval pass@1 and 39.1% in GSM8K exact match over the strongest evaluated quantization baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑