arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RSLM:用于近似最近邻搜索的无训练向量量化方法

RSLM: Training-Free Vector Quantization for Approximate Nearest Neighbor Search

Rastislav Lenhardt, Teodora Dobos, Thomas Vecchiato, Jiri Isa, Igor Ginzburg

arXiv 2608.30384首次发表:更新:

发表机构

Google; Technical University of Munich; University of Copenhagen(谷歌; 慕尼黑工业大学; 哥本哈根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无训练向量量化方法RSLM,将ANN搜索的嵌入压缩至每维度2-4比特,通过校正重构向量L2范数等创新,在降低系统复杂度与内存成本的同时保持或提升召回率。

AI 中文摘要

通过引入RSLM(旋转缩放劳埃德-麦克斯韦)——一类将嵌入压缩至每维度1-4比特的无训练向量量化编解码器,我们降低了典型大规模近似最近邻(ANN)搜索系统的内存成本与内存带宽,同时降低其复杂度,并在多个基准数据集上保持或提升了召回率。当前最先进的系统通过粗分区过滤候选、近似评分以缩小集合,再用更高精度表示(通常每维度≥8比特)对最优候选重评分。我们的相对编解码器可将此压缩至每维度2-4比特。我们利用ANN系统的特性,对近似评分阶段和重评分阶段均编码残差向量而非完整向量。由于最大内积搜索(MIPS)对向量范数非常敏感,我们校正量化向量的L2范数。我们的主要创新在于校正最终重构向量的L2范数,而非仅校正残差。这种重缩放替代了更复杂的方案,如各向异性损失。残差化方案使我们获得比通用量化方法更优的质量-规模权衡。我们的高性能实现利用了块级级联快速沃尔什-哈达玛变换(FWHT),其复杂度近似线性;采用AVX SIMD优化的码本;以及缩放因子的隐写式编码,以实现完美的缓存行对齐。

英文摘要

By introducing RSLM (Rotated Scaled Lloyd-Max), a family of training-free vector quantization codecs compressing embeddings to 1--4 bits per dimension, we reduce memory cost and memory bandwidth of a typical large-scale Approximate Nearest Neighbor (ANN) search system, while reducing its complexity and keeping or improving recall across multiple benchmark datasets. State-of-the-art systems filter candidates using coarse partitions, approximately score them to narrow the set, and then rescore the best with higher precision representations (often >=8 bits per dimension). Our relativized codecs can bring this down to 2--4 bits per dimension. We use the properties of the ANN system to encode residual vectors instead of full vectors, both for the approximate scoring phase and the rescoring phase. Since Maximum Inner Product Search (MIPS) is very sensitive to vector norms, we correct the $L_2$ norms of quantized vectors. Our major innovation is that we correct the $L_2$ norm of the final reconstructed vector rather than just the residual. Our rescaling replaces more complicated schemes, such as Anisotropic loss. The residualization scheme gives us a more favorable quality vs size trade-off than generic quantization methods. Our high-performance implementation leverages a block-wise cascaded Fast Walsh-Hadamard Transform (FWHT) with linear-like complexity, AVX SIMD-optimized codebooks, and a steganographic encoding of scaling factors for perfect cache-line alignment.

Comments14 Pages, 3 Figures, Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑