arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06564cs.LGcs.CL

量化损伤是乘法性的,而非加法性的

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

  • Holistic AI(霍利斯提克人工智能公司)
  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

Zekun Wu, Swati Dhiman, Adriano Koshiyama

中文总结 AI 辅助

该研究发现大型语言模型的量化损伤是乘法性的,而非加法性的,提出边际收缩概念,拟合关系可准确预测决策翻转率,增加1位是修复损伤最廉价的方式。

中文摘要 AI 辅助

量化是大型语言模型实际部署的方式,已知当量化位宽低于4位时会造成性能损伤,但没人能确定在给定位宽下模型的哪些决策会发生变化。这种损伤是隐蔽的:压缩后的智能体会停止调用工具,随后其安全弃权(不执行)的数量会减少一半,然而基准测试分数几乎没有变化。现有研究假设量化会添加大小大致固定的噪声,这会让置信度高的决策保持安全。我们转而直接测量决策本身:双向决策的边际是模型为其选择的选项得分减去其最佳替代选项的得分;我们在8个模型系列的16个模型、3种量化方法、位宽从8位降至2位的场景下,对比量化前后的边际变化。量化不会向边际添加固定大小的噪声,而是会将边际乘以一个随位宽下降的因子(4位时中位数为0.86,3位时为0.33,2位时为0.00),我们将此称为边际收缩。这种收缩降低了大边际所能提供的保护,模型自身的微小偏差会决定失败的方向:在3位时,调用工具的决策会收缩至不行动,而选择哪个工具的决策不受影响。在拟合统计比较中,加法噪声模型在受损的工具决策和安全决策上均未胜出。拟合的关系对保留的决策的翻转率的预测中位数误差为1.8个百分点,且拟合过程未使用任何翻转数据;对于每个决策,预测的翻转概率是校准后的不确定性估计(131758个预测的预期校准误差为0.004)。我们测量的所有模型都呈现出相同的关系,但常数是每个模型独有的,无法迁移。通过测量每个模型和位宽的小配对边际集,可在无需完整生成式评估的情况下估计哪些决策会失效;在我们的成本匹配测试中,没有任何方法比多增加1位能更廉价地修复损伤。

英文摘要

Quantization saves memory by storing model weights with fewer bits. It can also change model decisions, such as whether to call a tool or which option to choose from a finite set. We study these decision changes in 16 language models from 8 families at 4, 3 and 2 bits, across several post-training quantization settings. Our evaluation covers tool use, safety, general knowledge and social bias, using BFCL, XSTest, MMLU, BoolQ, BBQ and synthetic tasks. The decision margin is the score difference between two possible first tokens, measured before and after quantization. Writing the margin before quantization as $m$ and the margin after quantization as $m'$, we find an approximately linear relationship across decisions: $m' \approx c m + b$. The slope $c$ is usually below one and becomes smaller as precision falls, so quantization progressively shrinks decision margins. The offset $b$ is the same for every decision of one kind. Quantization therefore does not simply add random noise, and even a strong preference at full precision can flip. Quantization also affects different kinds of decisions to different degrees. Within tool use, whether to call a tool is often more sensitive than which tool to call: on 400 BFCL tasks, three of five models lose more completed calls than correct tool selections at 3-bit round-to-nearest. Under GPTQ and GGUF far fewer whether-to-call decisions flip than under plain rounding, so there is no single 3-bit failure point. The same relationship predicts how often decisions flip. Across 1,082 combinations of models, quantization settings, bit-widths and decision types, we fit the slope, the offset and the spread around the fitted line on half of the decisions and predict the flip rate on the other half. The predicted flip rate differs from the observed flip rate by a median of 1.0 percentage point, while reusing the flip rate of the first half misses by 1.3.

补充信息

↑