arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23125cs.LG

困惑度代价低估了激活量化破坏的内容

Perplexity Cost Understates What Activation Quantisation Breaks

Anish Sathyanarayanan

首次发表
浏览论文内容

中文总结 AI 辅助

本文发现困惑度作为聚合指标会掩盖激活量化对检索能力的严重破坏,提出旋转基量化可显著恢复归纳能力,但检索恢复有限。

中文摘要 AI 辅助

激活量化通常用一个聚合指标——困惑度来评估,该指标是对模型预测的每个 token 取平均。我们探究这个平均值是否能识别量化器损害了哪些计算。结果发现,困惑度是一个可靠的聚合信号:在来自四个家族的 12 个模型以及 780 次模型内比较中,困惑度偏好所保留的臂也分别在除 2.1% 和 4.0% 的情况外保留了更多的归纳和检索能力。但是,当困惑度仅上升了 1.2 到 1.5 倍时,归纳能力仍保持其完整准确率的 0.959,而检索能力已降至 0.554,这一差距是聚合数字所无法呈现的。这一差距具有结构性,而不仅仅是大小问题:一个匹配的、具有相同每通道幅值的高斯噪声对照实验几乎不会影响它,而仅随机化量化误差的符号(保持所有幅值不变)也几乎无害,因此幅值本身并不能解释这种损害。在旋转基中进行量化,改变了坐标对齐方式但不改变误差幅值,在单块干预中,以每 token 三个平均比特数将归纳能力从 0.001 恢复到 0.980,尽管在相同设置下检索能力的恢复不那么完全(0.694);在端到端四个平均比特数下,归纳能力达到 0.968,检索能力为 0.534。这种模式在另外两个高达 32B 参数的模型上以及在我们测试的部署配置中(在激活被推至 4 比特时,在 AWQ 下)均成立。困惑度目标限制了应用于激活的变换的平均代价;它本身并不能显示哪些计算得以幸存。

英文摘要

Activation quantisation is usually evaluated with an aggregate metric, perplexity, averaged over every token a model predicts. We ask whether that average identifies which computations a quantiser damages. Perplexity turns out to be a reliable aggregate signal: across 12 models from four families and 780 within-model comparisons, the arm perplexity prefers also retains more induction and more retrieval in all but 2.1 and 4.0 percent of cases respectively. But where perplexity has risen by only a factor of 1.2 to 1.5, induction still keeps 0.959 of its intact accuracy while retrieval has already fallen to 0.554, a gap the aggregate number does not surface. This gap has structure, not just size: a matched Gaussian-noise control of the same per-channel magnitude leaves it largely intact, and randomising only the sign of the quantisation error, every magnitude held fixed, is nearly as harmless, so magnitude alone does not explain the damage. Quantising in a rotated basis, which changes coordinate alignment without changing error magnitude, restores induction from 0.001 to 0.980 at three average bits per token in a single-block intervention, though retrieval recovers less completely at the same setting (0.694); end-to-end at four average bits, induction reaches 0.968 and retrieval 0.534. The pattern holds on two further models up to 32B parameters and, in the deployed configurations we tested, under AWQ once activations are pushed to 4 bits. A perplexity target bounds the average cost of a transformation applied to the activation; it does not, by itself, show which computations survived.

发表机构

  • BITS Pilani, K K Birla Goa Campus(比拉理工学院皮拉尼校区,K K Birla 果阿校区)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑