arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

免费保留物品语义:重新思考基于大语言模型的生成式推荐中的令牌初始化

Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation

Donald Loveland, Liam Collins, Bhuvesh Kumar, Danai Koutra, Neil Shah

arXiv 2608.07816首次发表:更新:

发表机构

Snap Inc.; University of Michigan(斯奈普公司; 密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对基于LLM的生成式推荐中SID令牌初始化的缺陷,提出质心初始化方法,无需额外开销即可提升推荐性能,减少训练步骤并改善冷物品表现。

AI 中文摘要

生成式推荐(GR)领域的最新进展利用大语言模型(LLM)作为推荐系统的主干,使LLM能够根据物品交互历史直接生成推荐。在这些系统中,物品通常通过语义ID(SIDs)表示,这些ID作为特殊令牌被添加到LLM的词汇表中。理想情况下,SIDs会赋予物品令牌表示语义先验,从而提升模型的泛化能力。然而,标准的词汇扩展通常将这些令牌初始化为随机高斯向量,丢弃了SIDs潜在的连续几何结构,迫使LLM从交互数据中重新学习令牌关系。为了证明这种设计的后果,我们首先表明,从该初始化开始训练往往会将SID嵌入围绕物品流行度而非语义进行组织。我们进一步表明,尽管持续预训练(CPT)在一定程度上降低了对流行度的依赖并改善了冷物品性能,但这种计算成本高昂的过程无法可靠地恢复原始语义几何结构。为解决这些发现,我们提出了一种简单、无参数的干预措施,直接从语义嵌入空间中对应的质心初始化SID令牌嵌入。这种即插即用方法仅需几行代码,无需额外的训练或推理开销,可将纯SFT的Recall@5提升多达16%,在减少多达40%的SFT步骤时达到峰值性能,并将冷物品的Recall@5提升多达60%。此外,在受益于额外CPT的数据集上,质心初始化可达到可比性能,同时仅需一半的CPT轮数。总体而言,我们的发现表明,保留SID几何结构(超越共享前缀结构)为基于LLM的GR提供了一种简单有效的语义先验。

英文摘要

Recent advances in generative recommendation (GR) leverage large language models (LLMs) as recommender backbones, enabling LLMs to directly generate recommendations conditioned on item-interaction histories. In these systems, items are often represented through semantic IDs (SIDs) added to the LLM vocabulary as special tokens. Ideally, SIDs imbue item token representations with semantic priors, thereby improving model generalization. However, standard vocabulary expansion typically initializes these tokens as random Gaussian vectors, discarding the SIDs' underlying continuous geometry and forcing the LLM to relearn token relationships from interaction data. To demonstrate the consequences of this design, we first show that training from this initialization tends to organize SID embeddings around item popularity rather than semantics. We further show that, despite partially reducing the reliance on popularity and improving cold item performance, the computationally expensive process of continual pretraining (CPT) fails to reliably recover the original semantic geometry. To address these findings, we propose a simple, parameter-free intervention that initializes SID token embeddings directly from their corresponding centroids in the semantic embedding space. Requiring only a few lines of code and no additional training or inference overhead, this drop-in approach improves pure-SFT Recall@5 by up to 16%, reaches peak performance with up to 40% fewer SFT steps, and improves cold-item Recall@5 by up to 60%. Moreover, on datasets that benefit from additional CPT, centroid initialization reaches comparable performance while requiring half as many CPT epochs. Together, our findings show that preserving SID geometry, beyond shared-prefix structure, provides a simple and effective semantic prior for LLM-based GR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑