令牌即所需:用于在推荐系统中实现大语言模型级I/O效率的双用途语义ID
Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems
浏览论文内容
中文总结 AI 辅助
研究针对大规模推荐系统的“内存墙”瓶颈,受计算机视觉数据压缩启发,提出双用途语义ID,用分层量化将连续嵌入压缩为离散语义ID执行协作身份和内容重建角色,经实验验证该方法能有效减少系统开销和数据占用,实现高效推荐。
中文摘要 AI 辅助
大规模推荐系统因大量密集的嵌入表而面临“内存墙”瓶颈。虽然生成式检索使用离散令牌作为ID,但高维上下文仍依赖低效的密集格式。受计算机视觉数据压缩启发,我们提出双用途语义ID以实现大语言模型级的I/O效率。我们的方法使用分层量化将连续嵌入压缩为离散语义ID,其执行两个并发角色:(1)协作身份:通过可学习的嵌入表对用户-项目交互进行建模;(2)内容重建:使用轻量级语义解码器进行实时嵌入近似。此方法用按需重建取代大量向量存储,减少系统开销和数据占用。我们通过离线评估和在一个主要视频共享平台的生产规模排名和检索系统中的成功在线部署证明了我们框架的有效性,表明离散令牌确实是高效、内容丰富的推荐所需的全部。
英文摘要
Large-scale recommendation systems face "Memory Wall" bottlenecks due to massive, dense embedding tables. While generative retrieval uses discrete tokens for IDs, high-dimensional context still relies on inefficient dense formats. Inspired by computer vision data compression, we propose Dual-purpose Semantic IDs to achieve LLM-level I/O efficiency. Our methodology uses hierarchical quantization to condense continuous embeddings into discrete Semantic IDs performing two concurrent roles: (1) Collaborative Identity: modeling user-item interactions via learnable embedding table; and (2) Content Reconstruction: using a lightweight Semantic Decoder for on-the-fly embedding approximation. This approach replaces massive vector storage with on-demand reconstruction, reducing system overhead and data footprints. We demonstrate the efficacy of our framework through offline evaluations and successful online deployment in production-scale ranking and retrieval systems at a major video sharing platform, showing that discrete tokens are indeed all you need for highly efficient, content-rich recommendation.
发表机构
- Google(谷歌)
- Google Deepmind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。