arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

说得出,却不知道:激进GGUF量化的小语言模型仍能写出它们已无法定义的罕见词

Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define

Saurabh Kumar Singh, Yogeshwar Singh Dadwhal, Malhar Vedak

arXiv 2610.04403首次发表:更新:

发表机构

Defence Institute of Advanced Technology(国防高级技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究审计27个GGUF量化小语言模型,发现激进量化至Q2级别时,子2B模型罕见词定义能力大幅下降,而生成流畅性保留,需按产物逐一验证。

AI 中文摘要

将开源权重语言模型量化至GGUF格式的混合精度K-quants,是使其在消费级硬件上运行的一种常见做法,但这一过程对模型细粒度词汇能力的影响尚未得到充分刻画。我们对13个模型家族、4种架构骨干、参数规模从0.35B到14B的27个量化产物进行了审计,沿其发布的量化阶梯评估至Q2_K(约每权重2.6比特),测试集包含429个经频率验证的罕见英语单词,采用两种探针:对提示词及其单句定义的表面包含性,通过分层多同义词匹配器评分,其误差通过盲法LLM裁判对所有定义进行普查并辅以人工验证来衡量。在Q2级别出现了三种情况:完全崩溃为不可用的构建、子2B模型中的严重语义分离,以及约3B以上模型的大体稳健保留。在每个子2B产物中,定义得分比基线低20%至67%,通常是包含性损失的数倍。两个对照组将罕见性与任务难度区分开来:在罕见词集合内,七个子2B产物中有六个的损失随罕见性增加而上升;在100个常见词集合上,所有八个产物中罕见词的损失均大于常见词,其中六个具有显著性。分词器词汇量大小不能预测损伤(Spearman rho=0.12);参数量起主导作用(rho=0.72),并在六个同分词器家族中的五个得到确认。Q4_K_M在参数规模≥1B时保持词汇上的清洁。这种损伤是频率分级的、依赖提供方的,并且不能通过WikiText-2困惑度来校准:在九个产物匹配的阶梯中,几乎相同的Q2惩罚(44.7%/47.6%)区分了一个保留其定义(3.6%)的产物与一个丢失定义(43.6%)的产物。激进量化的小模型可以继续生成流畅的文本,却不再知道其含义,这在语义具有后果的领域中给受硬件约束的部署带来风险。验证必须针对每个产物逐一进行。

英文摘要

Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, evaluated down their published ladder to Q2_K (about 2.6 bits per weight), on 429 frequency-validated rare English words under two probes: surface inclusion of a prompt-supplied word and its one-sentence definition, scored by a tiered multi-synonym matcher, its error measured by a blind LLM-judge census of every definition, with human verification. Three regimes emerge at Q2: total collapse into unusable builds, severe semantic dissociation in sub-2B models, and mostly robust preservation above about 3B. In every sub-2B artifact, definitions fall 20-67% below the artifact's baseline, typically several times the inclusion loss. Two controls separate rarity from task difficulty: within the rare set, loss rises with rarity in six of seven sub-2B artifacts, and on a 100-word common-word set rare words lose more than common words in all eight, significantly in six. Tokenizer vocabulary size does not predict the damage (Spearman rho=0.12); parameter count dominates (rho=0.72), confirmed within five of six same-tokenizer families. Q4_K_M remains lexically clean at >=1B. The damage is frequency-graded, provider-dependent, and not calibrated by WikiText-2 perplexity: across nine artifact-matched ladders, near-identical Q2 penalties (44.7%/47.6%) separate an artifact keeping its definitions (3.6%) from one losing them (43.6%). Aggressively quantized small models can keep generating fluent text while no longer knowing what it means, risking hardware-constrained deployments in domains where semantics carries consequences. Validation must be per artifact.

Comments41 pages, 15 figures, 1 table. Data, answer key, scored outputs and scorer code: https://github.com/NietzscheDostoevsky/snk-release

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑