arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化对检索的未察觉损害及其修复方法

The Undetected Damage of Quantization on Retrieval and How to Fix It

Luca Zhou, Alessandro Zirilli, Daniele Solombrino, Roberto Dessì, Emanuele Rodolà

arXiv 2609.24322首次发表:更新:

发表机构

Sapienza University of Rome; Paradigma; Not Diamond(罗马第一大学; Paradigma; Not Diamond)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现量化模型在保持分类准确率时仍改变14-46%的检索结果,提出利用最高两得分差距预测并修复量化损害,在检索中分配额外位宽、在分类中路由低差距输入至全精度。

AI 中文摘要

我们证明,一个保持分类准确率的量化模型仍会改变其14%至46%的top-1检索结果,且聚合排名指标仅能揭示部分此类损害。我们将此失败归因于最高两个得分之间的差距,并利用该差距来决定量化答案何时可被信任以及应在何处投入额外精度。我们证明,仅当该差距超过最大舍入误差的两倍时,top-1结果才能保证在量化后幸存。在分类中,得分是logits,损失函数将正确类别推离其他类别,从而鼓励这一差距。在检索中,得分是查询-文档得分,没有任何因素将top-1项与第二项分开。该差距无需标签即可测量。在部署前,它可预测哪些模型会在量化下失效;在部署时,它针对每个输入告知量化答案是否仍与全精度答案匹配。大多数分类输入具有足够宽的差距以信任量化答案,但很少有检索查询如此。该差距促使每项任务采用不同的修复方法。在检索中,将额外位宽花费在量化移动差距最大的层上,可以以一半的成本恢复额外位宽收益的四分之三。在分类中,将少数低差距输入路由至全精度,可以以一小部分成本恢复大部分丢失的准确率。

英文摘要

We show that a quantized model that keeps its classification accuracy still changes $14$ to $46\%$ of its top-1 retrieval results, and that aggregate ranking metrics reveal only part of this damage. We tie this failure to the gap between the two highest model scores and use that gap to decide when to trust a quantized answer and where additional precision should be spent. We show that the top-1 result is guaranteed to survive quantization only when this gap exceeds twice the largest rounding error. In classification, the scores are logits, and training compares the correct class against every other class, which encourages this gap. In retrieval, the scores are query-document similarities, and training compares each positive only against sampled negatives, so nothing separates the top-1 item from the second. This gap can be measured without labels. Before deployment, it predicts which models will break under quantization, and at deployment time it tells, per input, whether the quantized answer still matches the full-precision answer. Most classification inputs have a gap wide enough to trust the quantized answer, but few retrieval queries do. That gap motivates a different fix in each task. In retrieval, spending extra bit-width on the layers whose quantization moves the gap most recovers up to three-quarters of an extra bit's benefit for half its cost. In classification, routing the few low-gap inputs to full precision recovers most of the lost accuracy at a fraction of the cost.

Comments5 figures, 3 tables in the main paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑