arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化放大确定性而非偏见:服务时权重压缩的尺度依赖行为效应

Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression

Dachi Kurtskhalia

arXiv 2609.07901首次发表:更新:

AI 中文总结

研究发现量化压缩模型权重时,8B模型输出多样性下降(确定性增强),而14B/32B模型出现风格漂移,但均未发现偏见增强,表明量化主要放大确定性而非偏见。

AI 中文摘要

权重量化在很大程度上决定了开放权重大型语言模型服务的经济性。其成本通常通过能力基准来评估,在这些基准上,中等规模模型的4比特量化通常被认为是“几乎免费”的。我们研究了一个不同的问题:当多个答案都有效时,量化是否会改变模型选择说什么?我们在三种权重精度(W4A16 AWQ、W8A16 FP8-Marlin和bf16)下服务三个检查点(Qwen3-8B/14B/32B),保持硬件、软件和采样配置恒定,并收集了约71,000个按提示和种子配对的完成结果,跨越两个自定义的、经过泄漏检查的提示电池。我们在版本控制中预先指定了三波分析。在8B规模下,int4降低了输出多样性:同一场景下两个样本推荐同一品牌的概率增加了5.1个百分点(提示配对的符号翻转检验,Holm p=.023;在完整重新生成该臂时复现为+4.4个百分点),词汇多样性显著下降(TTR -0.011,标准化效应-0.51;对长度控制度量具有稳健性)。在14B和32B规模下,没有任何内容集中度度量达到显著性;相反,出现了风格漂移(14B时破折号率+0.46/千词,32B时+0.61/千词,两者Holm p≤.0024)。预先指定的刻板印象方向检验在每个尺度上均为零结果:输出集中在每个提示的模态答案上,而非刻板印象答案上。在机制上,词元级分布变得更平坦(决策词元熵+0.091比特,p=.015),而语义分布(直接从首词元对数概率测量)变得更集中(碰撞+2.6个百分点,p=.023):即使含义变得更重复,单个词元也变得更难预测。在8B(测试的最小规模)下,AWQ-int4服务可测量地缩小了建议的范围;审计应评估集中度以及偏见。

英文摘要

Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4-bit quantization of mid-sized models is often considered "nearly free." We examine a different question: when several answers are valid, does quantization change what a model chooses to say? We serve three checkpoints (Qwen3-8B/14B/32B) at three weight precisions (W4A16 AWQ, W8A16 FP8-Marlin, and bf16), holding the hardware, software, and sampling configuration constant, and collect approximately 71,000 completions paired by prompt and seed across two custom, leak-checked prompt batteries. We pre-specified the analyses in three waves in version control. At 8B, int4 reduces output diversity: the probability that two samples for the same scenario recommend the same brand increases by 5.1 percentage points (prompt-paired sign-flip test, Holm p = .023; reproduced at +4.4pp on a full regeneration of the arm), and lexical diversity falls substantially (TTR -0.011, standardized effect -0.51; robust to a length-controlled measure). At 14B and 32B, no content-concentration measure reaches significance; instead, stylistic drift emerges (em-dash rate +0.46/1k words at 14B and +0.61/1k at 32B, both Holm p <= .0024). Pre-specified tests of stereotype direction are null at every scale: outputs concentrate on the modal answer for each prompt rather than on stereotypical answers. Mechanistically, the token-level distribution becomes flatter (decision-token entropy +0.091 bits, p = .015) while the semantic distribution, measured directly from first-token log probabilities, becomes more concentrated (collision +2.6pp, p = .023): individual tokens become less predictable even as meanings become more repetitive. At 8B, the smallest size tested, AWQ-int4 serving measurably narrows the range of suggestions; audits should assess concentration as well as bias.

Comments8 pages, 3 figures. Under review at the NeurIPS 2026 Workshop on Deployable Small Foundation Models (LIGHT)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑