发表机构
Tsinghua University; Tencent Inc.; Shenzhen International Graduate School, Tsinghua University(清华大学; 腾讯公司; 清华大学深圳国际研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于流的物品分词器Tlow,解决传统推荐模型的参数过多与冷启动问题,经实验验证其能提升推荐性能,在微信多模态检索任务中使全局用户CTR提升10.32%、新物品提升11.64%。
AI 中文摘要
物品分词器将语义嵌入编码为令牌ID,以替代传统推荐模型中随机分配的物品ID,从根本上解决了参数过多和冷启动问题。然而,最常用的分词器RQ-VAE因其码本之间固有的依赖关系而存在解码效率低的问题。同时,诸如优化乘积量化(OPQ)之类的高效独立分词器仍难以应对语义嵌入的维度相关性和分布复杂性。在这项工作中,我们提出了一种基于流的物品分词器(Tlow),用于将原始语义嵌入转换为潜在空间,其中嵌入符合统一的标准正态分布,实现了维度独立性和分布简单性的双重优势。对这些潜在嵌入执行独立分词可产生语义清晰的令牌ID。此外,我们引入了一种新颖的码本引导机制,以使码本空间与令牌嵌入空间对齐,进一步助力学习更具语义区分度的令牌嵌入。在四个公共数据集上进行的离线实验表明,Tlow的分词和码本引导显著提升了推荐性能。在跨域和多模态推荐上的改进也证明了在简化的嵌入空间中物品分词的有效性。在中国最大的社交媒体平台微信上针对多模态检索任务的在线实验验证了Tlow强大的分布转换能力。基于令牌ID的检索模型在全局范围内将用户点击率(CTR)提升了10.32%,在新物品上提升了11.64%。我们的代码可在该https URL获取。
英文摘要
Item tokenizer encodes semantic embeddings into token IDs to replace the randomly assigned item IDs used in traditional recommendation models, fundamentally addressing the problems of excessive parameters and cold starts. However, the most common tokenizer, RQ-VAE, suffers from low decoding efficiency due to the inherent dependencies among its codebooks. Meanwhile, efficient independent tokenizers such as optimized product quantization (OPQ) still struggle with dimensional correlations and distribution complexity of semantic embeddings. In this work, we propose a f\underline{low}-based item \underline{T}okenizer (Tlow) to transform raw semantic embeddings into a latent space where embeddings conform to a unified standard normal distribution, achieving dual advantages of dimensional independence and distributional simplicity. Independent tokenization performed on these latent embeddings yields semantically clear token IDs. Additionally, we introduce a novel codebook guidance to align the codebook space with the token embedding space, further aiding the learning of more semantically distinct token embeddings. Offline experiments on four public datasets demonstrate that Tlow's tokenization and codebook guidance significantly improve recommendation performance. The improvement on cross-domain and multi-modal recommendations also proves the effectiveness of item tokenization in a simplified embedding space. Online experiments for a multi-modal retrieval task on China's largest social media platform WeChat validate Tlow's powerful distribution transformation capability. The retrieval model based on token IDs improves user CTR by 10.32\% globally and by 11.64\% for new items. Our codes are available at https://github.com/wjjln/Tlow.
CommentsCIKM'26 Applied Research