发表机构
ThakiCloud(ThakiCloud)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出无标签的混合精度训练后量化方法,利用量化引起的表示漂移作为敏感性信号,在文本嵌入器上实现高效分配,避免灾难性失败,但细粒度优化仍有待解决。
AI 中文摘要
混合精度训练后量化需要每个模块的敏感性信号;对于文本嵌入器而言,显而易见的信号——即模块被量化时检索质量所损失的程度——需要相关性标签,而部署环境很少具备这些标签。我们衡量了一种无标签的替代指标:量化引起的表示漂移,通过量化一个模块、重新编码语料库,并记录输出嵌入相对于其全精度位置的移动距离来获得。其特殊性在于可观测对象:密集检索器用于排序的部署输出表示。在五个开发嵌入器上,配置级漂移以宏观斯皮尔曼相关系数0.911对采样混合精度计划与保留的检索质量进行排序;在可用范围内,敏感性可跨校准语料库和检索领域迁移;模块漂移在排序上保持一致但数值上不可加;而基于相关性的敏感性并未增加一致价值。该方法是在硬性打包字节预算下的加法分配,无需标签且无需搜索。在三个直到方法、基线和假设被冻结并密封前未触碰的嵌入器上,预先注册的方向性假设相对于先前LieQ标准成立(在主预算下3/3,无崩溃),且漂移分数在2/3的情况下高于双侧LieQ强化基线;但在主预算下,漂移在数值上低于所有三个嵌入器的同预算均匀精度(-0.99、-0.85、-1.01个百分点),按设计减少了模块和整体模型漂移。因此,输出漂移是一种稳健的粗粒度敏感性信号,而非普遍最优的分配目标:它避免了迁移符号几何适配的灾难性失败,并可在均匀精度崩溃的紧张预算下保持可用,但在强均匀操作点周围的细粒度重新分配仍未解决。
英文摘要
Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.
Comments26 pages, 22 tables, 4 figures