arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QTEA:带稀疏残差显著权重与按列优化的三值大语言模型

QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi

arXiv 2609.00224首次发表:更新:

发表机构

University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

QTEA是一种亚2比特的三值后训练量化框架,通过按列优化、半结构化稀疏性和误差衰减提升三值LLM的精度与推理速度,在Qwen3-14B和Llama3-8B上均取得显著性能提升。

AI 中文摘要

仅权值的后训练量化(PTQ)可缓解大规模部署大语言模型(LLM)的计算负担,但现有PTQ方法常无法跨模型泛化,且在低于2比特时会出现严重精度损失。许多方法利用非结构化稀疏性缓解该损失,但会牺牲规则性和GPU友好型执行。我们提出QTEA,一种亚2比特的PTQ框架,将权重量化为三值,并以显著权重作为残差误差补偿器。为保持硬件效率,残差被分配至选定列,在显著列内采用半结构化1:4稀疏性。我们进一步在GPTQ风格的按列量化中添加按列缩放优化,交替更新每列的缩放因子和三值分配以降低重构误差。我们还识别出GPTQ中存在依赖顺序的误差传播,并引入误差衰减以缓解后期误差累积。在Qwen3-14B上,QTEA将所有权重压缩至每权重有效1.7比特,同时较最强的三值PTQ基线提升了16.7%的平均精度;在WikiText和C4上的困惑度分别降低1.40倍和2.61倍。该趋势在Llama3-8B上同样成立,QTEA在相同数据集上获得6.6%的精度提升和1.34倍/1.95倍的困惑度降低。最后,我们开发了基于查找表的内核,较FP16基线实现了7.2倍的每token生成速度提升。代码可在此httpsURL获取。

英文摘要

Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2 bits. Many leverage unstructured sparsity to mitigate this loss, but at the cost of regularity and GPU-friendly execution. We present QTEA, a sub-2-bit PTQ framework that quantizes weights into ternary values and uses salient weights as residual error compensators. To maintain hardware efficiency, residuals are assigned to selected columns with semi-structured $1:4$ sparsity within the salient columns. We further add column-wise rescale refinement to GPTQ-style column-by-column quantization, alternately updating per-column scales and ternary assignments to reduce reconstruction error. We also identify order-dependent error propagation in GPTQ and introduce error decay to attenuate late-stage error accumulation. On Qwen3-14B, QTEA compresses all weights to an effective 1.7 bits per weight while improving average accuracy over the strongest ternary PTQ baseline by 16.7%. It also achieves 1.40$\times$ and 2.61$\times$ lower perplexity on WikiText and C4 respectively. This trend holds on Llama3-8B, where QTEA obtains a 6.6% accuracy gain and 1.34$\times$ / 1.95$\times$ lower perplexity on the same datasets. Finally, we develop a lookup-table based kernel that achieves 7.2$\times$ faster per-token generation over an FP16 baseline. Code is available at https://github.com/Intelligent-Microsystems-Lab/QTEA.

CommentsAccepted by EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑