什么造就了生成式推荐的良好语义ID?一项可复现性研究
What Makes a Good Semantic ID for Generative Recommendation? A Reproducibility Study
- Shandong University(山东大学)
- University of Glasgow(格拉斯哥大学)
- Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过大规模可复现实验,系统探究生成式推荐中语义ID的设计因素,发现其效果非单调,无普遍最优设计,不同设计具有互补优势,为未来系统提供实用指导。
AI中文摘要:
生成式推荐已成为一个活跃的研究方向,其中物品通常由语义ID(SIDs)表示:逐token生成的离散编码。尽管实证结果强劲,但SID设计在构建策略、码本组织和编码长度上差异很大,使其对推荐性能的真实影响尚不明确。我们在统一实验框架下开展了一项大规模可复现性研究,以系统探究语义ID设计对生成式推荐的影响。我们聚焦于一个基本问题:什么造就了生成式推荐的良好语义ID?为回答此问题,我们考察了四个方面:不同语义ID设计的相对有效性、码本利用率与推荐质量之间的联系、语义编码长度的影响,以及语义ID设计对局部物品语义保持的作用。通过统一评估和额外的跨数据集受控分析,我们发现SID设计的效果在很大程度上是非单调的:没有一种SID设计普遍最优,常用的基于RQ-VAE和OPQ的设计在不同数据集上可能表现不一致。具有最均衡一级码本的方法并非始终是最佳推荐器,这表明利用率具有诊断性但不足以决定性能。扩展生成骨干网络或SID长度也并非总是有益的。最后,语义邻域分析揭示,没有一种SID设计在所有局部语义保持概念上占优;相反,不同设计展现出互补优势,这些优势在不同数据集和邻域大小上保持稳定。我们的研究提供了对语义ID设计的受控且可复现的理解,并为未来生成式推荐系统提供了实用见解。
英文摘要:
Generative recommendation has emerged as an active research direction, where items are commonly represented by semantic IDs (SIDs): discrete codes generated token by token. Despite strong empirical results, SID designs vary widely in construction strategy, codebook organization, and code length, making their true impact on recommendation performance unclear. We conduct a large-scale reproducibility study to systematically investigate the impact of semantic ID design on generative recommendation under a unified experimental framework. We focus on a fundamental question: What makes a good semantic ID for generative recommendation? To answer this question, we examine four aspects: the relative effectiveness of different semantic ID designs, the connection between codebook utilization and recommendation quality, the effect of semantic code length, and the influence of semantic ID design on local item semantic preservation. Through a unified evaluation and additional cross-dataset controlled analyses, we find that the effects of SID design are largely non-monotonic: no single SID design is universally best, and commonly used RQ-VAE- and OPQ-based designs can behave inconsistently across datasets. The method with the most balanced first-level codebook is not consistently the best recommender, showing that utilization is diagnostic but insufficient. Scaling either the generative backbone or the SID length is also not always beneficial. Finally, semantic-neighborhood analysis reveals that no single SID design dominates all notions of local semantic preservation; instead, different designs exhibit complementary strengths that remain stable across datasets and neighborhood sizes. Our study provides a controlled and reproducible understanding of semantic ID design and offers practical insights for future generative recommender systems.