用于生成式推荐的拓扑感知词元化
Topology-Aware Tokenization for Generative Recommendation
浏览论文内容
中文总结 AI 辅助
研究针对生成式推荐中项目词元化的拓扑失真问题,提出拓扑感知词元化框架TopoTok,通过多级蒸馏方案逐步恢复拓扑,实验表明该方法有效减轻失真,优于现有词元化器,提升了生成式推荐性能。
中文摘要 AI 辅助
生成式推荐将序列推荐重新表述为自回归生成任务,但该范式中的一个关键问题被忽视:项目词元化中的拓扑失真。具体而言,预训练语义嵌入空间中项目的内在邻接关系在量化后被显著破坏。这种拓扑失真误导了模型对项目相似性的感知,最终限制了生成式推荐的准确性。为解决此问题,我们提出了拓扑感知词元化(TopoTok),这是一个在整个量化层次结构中保留项目关系结构的项目词元化框架。与之前词元化中的整体监督不同,TopoTok引入了一种多级蒸馏方案,从粗粒度到细粒度逐步恢复拓扑:1)组间蒸馏以捕获全局聚类关系;2)组内蒸馏以细化语义聚类内的局部结构;3)项目间蒸馏以在单个项目级别强制进行细粒度对齐。在三个基准数据集上的大量实验表明,TopoTok有效地减轻了拓扑失真,始终优于现有词元化器,在Recall@5中实现了高达9.42%的显著性能提升。
英文摘要
Generative recommendation reformulates sequential recommendation as an autoregressive generation task, yet a critical issue in this paradigm remains overlooked: topology distortion in item tokenization. In particular, we observe that the intrinsic adjacency relationships of items in the pretrained semantic embedding space are significantly disrupted after quantization. This topology distortion misleads the model's perception of item similarity, ultimately bottlenecking the accuracy of generative recommendations. To address this issue, we propose Topology-Aware Tokenization (TopoTok), an item tokenization framework that preserves item relational structure throughout the quantization hierarchy. Different from the prior monolithic supervision in tokenization, TopoTok introduces a multi-level distillation scheme to progressively recover the topology from coarse to fine granularity: 1) Inter-Group Distillation to capture global cluster-wise relations; 2) Intra-Group Distillation to refine local structures within semantic clusters; and 3) Inter-Item Distillation to enforce fine-grained alignment at the individual item level. Extensive experiments on three benchmark datasets demonstrate that TopoTok effectively alleviates topology distortion and consistently outperforms state-of-the-art tokenizers, achieving significant performance gains of up to 9.42% in Recall@5.