arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PRQ-KMeans:用于语义ID分词的投影残差量化

PRQ-KMeans: Projection Residual Quantization for Semantic ID Tokenization

Yunxiao Luo, Siyuan Wang, Ben Chen, Chenyi Lei, Qingpeng Cai

arXiv 2608.24207首次发表:更新:

发表机构

Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语义ID分词的残差量化局限,提出PRQ-KMeans方法,在工业搜索及推荐基准上实现分词器最优性能,获命中率、平均倒数排名的显著提升。

AI 中文摘要

语义标识符(SIDs)将实体表示为分层令牌序列,用于生成式检索与推荐。残差量化分词器通过在每一层选择一个码字并将残差传递至下一层来构建这些序列。我们将该过程视为渐进式共性移除:每个令牌捕获其组内共享的组件,后续令牌则需建模剩余差异。此视角揭示了三个局限:语料库范围的共享组件会占用首层容量,硬分配忽略与邻近码字的分级相似度,全码字减法可能在后续残差中留下所选码字方向上的变异。因此,我们在事后设置中开发解决方案,残差构建不受输入重构约束。具体而言,我们提出PRQ-KMeans,其移除全局均值组件、通过Top-k相似度加权更新优化质心,并以投影残差替代全码字减法,该残差会移除每个表示的所选质心组件。在大规模工业搜索数据集及四个公开推荐基准上的实验显示,PRQ-KMeans在所有评估的分词器中实现了最强的整体性能,在工业数据集上的命中率(HitRate)提升最高达7.4%,平均倒数排名(MRR)提升最高达11.8%。

英文摘要

Semantic identifiers (SIDs) represent entities as hierarchical token sequences for generative retrieval and recommendation. Residual-quantization tokenizers construct these sequences by selecting a codeword at each level and passing a residual to the next. We view this process as progressive commonality removal: each token captures a component shared within its group, while later tokens should model the remaining differences. This view reveals three limitations: a corpus-wide shared component can consume first-level capacity, hard assignment ignores graded similarities to nearby codewords, and full-codeword subtraction can leave variation along the selected-codeword direction in the next residual. We therefore develop our solution in the post-hoc setting, where residual construction is not constrained by input reconstruction. Specifically, we propose PRQ-KMeans, which removes the global-mean component, refines centroids with Top-k similarity-weighted updates, and replaces full-codeword subtraction with a projection residual that removes each representation's selected-centroid component. Experiments on a large-scale industrial search dataset and four public recommendation benchmarks show that PRQ-KMeans achieves the strongest overall performance among the evaluated tokenizers, including gains of up to 7.4% in HitRate and 11.8% in MRR on the industrial dataset.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑