arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16391cs.IRcs.CL

训练后量化在何处破坏文本嵌入器:四个嵌入器家族的系统性测量图谱

Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families

Hyojung Han

首次发表
浏览论文内容

中文总结 AI 辅助

本研究系统测试了训练后量化在检索嵌入器上的传统建议,发现嵌入表保护、模块敏感性排序等均失效,并揭示蒸馏小模型在极端量化下可超越大模型。

中文摘要 AI 辅助

仅权重的训练后量化是缩小检索嵌入器最廉价的方式,而应用它的公认建议——保护嵌入表、按模块敏感性分配比特、优先选择排序感知目标而非权重重建——在很大程度上被原封不动地引入了大语言模型量化中。我们直接在检索嵌入器上测试该建议,对来自四个架构家族的五个检查点在比特宽度和组大小的网格上进行量化,并在每个宽度下分别隔离嵌入、注意力与前馈模块。每条启发式规则都未能如所述般迁移。嵌入表在任何家族中都未成为主导的隔离保护优先级,尽管在多个家族中它是最大的张量。模块敏感性无法作为可迁移的排序而存在:在INT4/g16下,模块间的差异过小以至于无法据此分配;在INT3下,排序变得依赖家族,且联合损伤不再是各部分之和;在INT2下,相当的重建误差伴随着从全精度的1.3%到65.9%不等的保留率。一个廉价的重建代理对于筛选均匀比特宽度是有用的,但在选择要保护哪些张量时可靠性显著降低;其在整个网格上表现出的强度是一种范围扩展的伪影。一个蒸馏的109M学生模型在INT3下以68.4 MB的尺寸保持了78.04的NDCG@10,并在尺寸和质量上均超越了其0.6B教师模型的极端PTQ分支(297.9 MB,64.46)——但仅限于其被蒸馏的任务内。尺寸是实际存在文件的字节数,而非算术估算,测量仓库为每一个文件携带字节来源。

英文摘要

Weight-only post-training quantization is the cheapest way to shrink a retrieval embedder, and the received advice for applying it -- protect the embedding table, allocate bits by module sensitivity, prefer a ranking-aware objective over weight reconstruction -- was carried into LLM quantization largely intact. We test that advice on retrieval embedders directly, quantizing five checkpoints from four architecture families across a grid of bit widths and group sizes, and isolating the embedding, attention and feed-forward blocks at each width. Every heuristic fails to transfer as stated. The embedding table never emerges as the dominant isolated protection priority in any family, despite being the largest tensor in several of them. Module sensitivity does not survive as a transferable ordering: at INT4/g16 the spread between modules is too small to allocate against, at INT3 the ordering becomes family-dependent and joint damage stops being the sum of its parts, and at INT2 comparable reconstruction error accompanies retention ranging from 1.3 to 65.9 percent of full precision. A cheap reconstruction proxy is useful for screening uniform bit widths but substantially less reliable for choosing which tensors to protect; its apparent strength across the whole grid is a range-extension artifact. A distilled 109M student at INT3 holds 78.04 NDCG@10 in 68.4 MB and dominates the extreme-PTQ arm of its own 0.6B teacher, 297.9 MB at 64.46, on both size and quality -- but only inside the task it was distilled for. Sizes are byte counts of files that exist rather than arithmetic estimates, and the measurement repository carries the byte provenance for every one of them.

发表机构

  • ThakiCloud(塔基云)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑