ARCHead:大语言模型输出头的激活度量残差修正
ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
浏览论文内容
中文总结 AI 辅助
ARCHead是一种LM-head压缩器,通过量化低秩核心、分组INT4残差及激活衍生低秩修正,将LM-head存储降3.7-3.9倍,在Qwen3-8B-Base上性能优于朴素INT4,可补充块量化器。
中文摘要 AI 辅助
仅权重量化可大幅减少大语言模型(LLM)Transformer块的存储需求,但实际后端常将最终语言建模头(LM-head)保留为BF16或FP16格式。直接量化该投影层会严重扰动词汇对数分布。本文提出ARCHead,一种打包式LM-head压缩器,结合量化低秩核心、分组INT4残差及基于激活衍生度量拟合的低秩修正。ARCHead不存储稠密BF16头,可将持久LM-head存储量降低3.7-3.9倍。在Qwen3-8B-Base上,其使用25.6%的BF16头存储量,相对困惑度达1.007;存储量匹配的朴素INT4方法相对困惑度为1.14-1.16。替换AWQ或bitsandbytes留下的BF16头仅增加0.006-0.007交叉熵,吞吐量变化小于2%。ARCHead可补充块量化器,压缩其未处理的大型输出投影层,代码可获取于该https URL。
英文摘要
Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection naively can strongly perturb the vocabulary-logit distribution. We present ARCHead, a packed LM-head compressor that combines a quantized low-rank core, group-wise INT4 residuals, and a low-rank correction fitted in an activation-derived metric. ARCHead stores no dense BF16 head and reduces persistent LM-head storage by 3.7-3.9x. On Qwen3-8B-Base, it uses 25.6% of BF16 head storage while attaining 1.007 relative perplexity; storage-matched naive INT4 yields 1.14-1.16. Replacing the BF16 head left by AWQ or bitsandbytes adds only 0.006-0.007 cross-entropy, with less than 2% throughput change in our measurements. ARCHead therefore complements block quantizers by compressing the large output projection they can leave untouched. Code is available at https://github.com/suayptalha/archead.
发表机构
- aiXplain, Inc.(aiXplain公司)
机构由 AI 辅助整理,请以论文原文为准。