发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型量化中的偏差问题,提出QuantiBias基准测试,通过结合多种控制和对比有无推理的构建来评估偏差,发现量化器按无偏差预防信号的数据分配精度,推理对部分偏差有减半效果,强调需重新评估量化模型的开放式偏差。
AI 中文摘要
几乎所有面向大众的大型语言模型都会进行量化:先以全精度训练,然后为提高效率进行压缩。这一步骤通常被认为无害且很少重新检查安全性。我们发现其主要副作用是标准安全评估遗漏的偏差增加。在模型、训练和提示固定的情况下,量化模型在拒绝有害请求、避免过度拒绝良性提示以及选择无偏差的多项选择题答案方面表现良好。然而,在回答开放式问题时,我们研究的所有八种语言中,该模型都会出现刻板印象,在独立评判下,约四分之一的开放式答案存在偏差(在整个压缩等级中为24%至27%):它通过了所有标准检查,但用户所接触到的偏差却明显增加。这种选择性差距是一个可靠的发现;开放式偏差是否会随着压缩进一步增加尚不确定,这对评判其得分很敏感。我们使用QuantiBias来解决这两个问题,这是一个基准测试,它将生成式多语言刻板印象探针与拒绝和多项选择控制相结合,以隔离开放式生成,对比有无推理的每种构建,并对生成内容的严重程度进行评级。在两个骨干模型(Qwen和Gemma)、五个系列筛选和八个基准测试中,量化器根据不携带偏差预防信号的能力数据分配额外精度,回答前的推理在某些系列中可将影响减半,而在其他系列中则无效。量化构建必须重新评估开放式偏差,而不仅仅是基于它已经通过的简短形式保护措施。
英文摘要
Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its principal side effect is increased bias that standard safety evaluation misses. Holding the model, its training, and the prompts fixed, a quantized model still refuses harmful requests, still avoids over-refusing benign prompts, and still selects the unbiased multiple-choice answer. Yet asked an open-ended question, the same model volunteers stereotypes in all eight languages we probe, in roughly one in four open-ended answers under an independent judge (~24% to ~27% across the compression ladder): it passes every standard check and still reaches users measurably more biased. The selective gap is a robust finding; whether open-ended bias further increases with compression is less certain, sensitive to the judge that scores it. We address both with \textbf{QuantiBias}, a benchmark that pairs a generative, multilingual stereotype probe with the refusal and multiple-choice controls that isolate open-ended generation, contrasts each build with and without reasoning, and rates the content severity of what it generates. Across two backbone models (Qwen and Gemma), a five-family screen, and eight benchmarks, quantizers allocate their extra precision by capability data that carries no bias-prevention signal, and reasoning before answering roughly halves the effect on some families while doing nothing on others. A quantized build must be re-evaluated for open-ended bias, not only on the short-form safeguards it already passes.
CommentsBenchmark protocol on Hugging Face: https://huggingface.co/datasets/emilioferrara/quantibias