AI 中文总结
本研究针对现有量化方案均匀分配比特的局限,提出可变比特分配框架,在相同存储预算下,对具套娃属性的嵌入,其使乘积量化与标量量化的召回率最高分别提升8%、18%,为大规模检索系统提供新方向。
AI 中文摘要
量化是处理现代模型生成的嵌入规模不断增长的基础技术。现有量化方案大多与嵌入无关,且在各维度均匀分配比特。然而,近期模型生成的嵌入具有显著的几何结构。本研究探讨在固定内存预算下,可变比特分配方案是否可提升量化质量。我们提出一种简单的可变比特分配框架,将嵌入划分为连续桶并在桶间非均匀分配存储,采用贪心分配策略,将该框架实例化为乘积量化(Product Quantization, PQ)与标量量化(Scalar Quantization, SQ)两种形式。我们对已知具有套娃属性(Matryoshka property, MRL)的嵌入开展一系列实验,在相同存储预算下,非均匀分配始终优于均匀基线。最大提升出现在低比特 regime,此时均匀分配对 MRL 嵌入效率极低。在相同压缩率下,可变分配使 PQ 的召回率提升最高达 8%,SQ 的召回率提升最高达 18%。我们的结果为大规模检索系统的结构感知压缩与索引技术指明了新方向。
英文摘要
Quantization is a fundamental technique to handle the growing sizes of embeddings generated by modern models. Existing quantization schemes are largely embedding agnostic and allocate bits uniformly across dimensions. However, recent models produce embeddings with significant geometric structure. In this work, we investigate whether a variable bit allocation scheme can improve quantization quality under a fixed memory budget. We propose a simple variable bit allocation framework that partitions an embedding into contiguous buckets and allocates storage non-uniformly across them. Using a greedy allocation strategy, we instantiate this framework for both Product Quantization (PQ) and Scalar Quantization (SQ). We perform a series of experiments on embeddings known to have the Matryoshka property (MRL), and consistently observe that non-uniform allocations outperform uniform baselines at identical storage budgets. The largest improvements occur in the low-bit regime, where uniform allocation is particularly inefficient for MRL embeddings. At the same compression rates, variable allocation improves recall by up to 8\% for PQ and up to 18\% for SQ. Our results suggest a new direction for structure-aware compression and indexing techniques for large-scale retrieval systems.
CommentsAccepted at the 2nd Workshop on Vector Databases (VecDB), part of 52nd International Conference on Very Large Data Bases (VLDB 2026)