发表机构
Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对生成式推荐中多级残差量化的缺陷,提出单级大语义码本及动态更新机制,在公开数据集、服务架构及在线测试中均实现推荐性能与效率的提升。
AI 中文摘要
生成式推荐用离散语义ID(SIDs)序列表征每个物品,并预测该序列以检索下一个物品。典型系统采用多级残差量化,这会增加自回归解码成本,且会产生可能稀疏占用的大型分层空间。随着新物品到来和曝光分布变化,静态码本也会与当前流量失配。我们提出一种单级大语义码本,用一个语义 token 替代多个残差语义码,同时保留单独的协作消歧 token 以减少物品碰撞。我们进一步引入基于时间权重衰减、指数移动平均中心更新以及SID变化的曝光加权惩罚的感知曝光动态更新机制。我们还开发了涵盖表征质量、码本利用率、集群负载、全SID碰撞和时间稳定性的离线评估框架。在两个公开数据集上,两级SID使OneRec-V1的平均Recall@10提升5.0%-8.8%、平均NDCG@10提升4.1%-5.1%,使OneRec-V2的对应指标分别提升7.1%-8.7%、3.8%-8.5%。动态更新在KuaiRec上进一步提升效果。在三种服务架构下,更短的SID将估计的自回归解码FLOPs降低47.93%-48.70%,并使单卡QPS提升28.57%-47.0%。针对2.5%生产流量的五天在线A/B测试,使主要消费指标提升0.792%。
英文摘要
Generative recommendation represents each item with a sequence of discrete Semantic IDs (SIDs) and predicts the sequence to retrieve the next item. Typical systems use multi-level residual quantization, which increases autoregressive decoding cost and creates a large hierarchical space that may be sparsely occupied. Static codebooks also become misaligned with current traffic as new items arrive and exposure distributions change. We propose a single-level large semantic codebook that replaces multiple residual semantic codes with one semantic token while retaining a separate collaborative disambiguation token to reduce item collisions. We further introduce an exposure-aware dynamic update mechanism based on temporal weight decay, exponential moving-average center updates, and an exposure-weighted penalty on SID changes. We also develop an offline evaluation framework covering representation quality, code utilization, cluster load, full-SID collision, and temporal stability. On two public datasets, the two-level SID improves mean Recall@10 by 5.0%-8.8% and mean NDCG@10 by 4.1%-5.1% for OneRec-V1, and by 7.1%-8.7% and 3.8%-8.5%, respectively, for OneRec-V2. Dynamic updating provides further gains on KuaiRec. Across three serving architectures, the shorter SID reduces estimated autoregressive-decoding FLOPs by 47.93%-48.70% and increases single-card QPS by 28.57%-47.0%. A five-day online A/B test serving 2.5% of production traffic improves the primary consumption metric by 0.792%.
Comments6 figures, 10 tables, and 1 algorithm