arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言条件去量化:从非英语语言中恢复量化所损失的能力

Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages

Nirmal Thomas

arXiv 2608.11786首次发表:更新:

发表机构

Prathama International(普拉塔玛国际)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对激进量化导致非英语语言能力受损的问题,提出语言条件去量化(LCD)方法,在参数量级低于40亿的模型上有效恢复非英语语言的困惑度与准确率差距,性能优于同类基线方法。

AI 中文摘要

激进的量化会对多语言能力造成不成比例的损害:在参数量级低于40亿的INT3 GPTQ模型中,我们测得非英语语言的困惑度下降幅度是英语的2-4倍。我们提出了语言条件去量化(LCD),这是一种事后方法,它为已量化模型的线性层附加每语言的秩2 LoRA修正,每语言仅增加0.12%的参数,且在单GPU上训练时间不到20分钟。在Qwen2.5-3B和Llama-3.2-3B模型上,LCD恢复了非拉丁文字语言70-83%的困惑度差距,以及GlobalMMLU准确率差距的17-28%,在类型学差异较大的语言上,其表现优于同等容量的语言无关修正方法3-9个百分点,且比无数据低秩基线(LQER)高出一个数量级。我们进一步发现了困惑度与准确率之间的脱节现象,并将其归因于量化损伤的集中位置:Llama模型的早期层错误会向下游传播且难以被局部修正,而Qwen模型的后期层错误则不会,LCD的层受限变体直接验证了这一机制。

英文摘要

Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on English. We propose Language-Conditional Dequantization (LCD), a post-hoc method that attaches per-language rank-2 LoRA corrections to the linear layers of an already-quantized model, adding 0.12% parameters per language and training in under 20 minutes on a single GPU. Across Qwen2.5-3B and Llama-3.2-3B, LCD recovers 70-83% of the perplexity gap for non-Latin script languages and 17-28% of the GlobalMMLU accuracy gap, outperforming a language-agnostic correction of equal capacity by 3-9 points on typologically distant languages and a data-free low-rank baseline (LQER) by an order of magnitude. We further identify a perplexity-accuracy disconnect and trace it to where quantization concentrates damage: early-depth errors (Llama) propagate downstream and resist local correction, while late-depth errors (Qwen) do not. A layer-restricted variant of LCD validates this mechanism directly.

Comments9 pages, 1 figure, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑