arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36670cs.AI

FineSID:面向生成式推荐的可扩展且高效的语义标识符学习

FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation

Song-Li Wu, Weinan Gan, Zhaocheng Du, Xianquan Wang, Jingyi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对生成式推荐中语义标识符学习因Top-1分配导致梯度稀疏和码本冲突的问题,提出FineSID统一量化框架,通过软性全码本梯度传播实现平衡优化,提升码本利用率和推荐准确性。

中文摘要 AI 辅助

生成式推荐的一个关键先决条件是设计语义标识符(SIDs),这些标识符既要能扩展到大型项目集,又要能被高效地学习。现有的SID学习方法从根本上依赖于向量量化过程中的Top-1硬分配。虽然启发式策略(如基于聚类的初始化或强制的事后冲突解决)可以人为地提高码本覆盖率,但它们往往会破坏端到端的语义对齐,并且无法解决潜在的优化瓶颈:稀疏梯度传播。在标准的Top-1分配中,梯度集中在频繁选中的码字的一个狭窄子集上,导致大多数码字本质上训练不足,并引发严重的SID冲突。为了在不依赖复杂初始化先验的情况下从本质上克服这一限制,我们提出了FineSID,一个统一的量化框架,它超越了Top-1分配,能够在整个码本上实现细粒度的梯度传播。FineSID不是只更新一个选中的码字,而是以软性、可微分的方式将学习信号分配给所有码字。这种设计促进了全局平衡的码本优化,同时严格保持语义一致性,有效缓解了SID冲突,并在大规模、高维码本中稳定了训练。在多个公共基准上的大量实验表明,FineSID对初始化配置具有鲁棒性,并且能持续提高码本利用率和推荐准确性。我们的工作为语义标识符学习提供了一个有原则的、与初始化无关的解决方案,推进了生成式推荐的实用性。

英文摘要

A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propagation. In standard Top-1 assignment, gradients concentrate on a narrow subset of frequently selected codewords, leaving the majority inherently under-trained and causing severe SID collisions. To overcome this limitation natively without relying on complex initialization priors, we propose FineSID, a unified quantization framework that moves beyond Top-1 assignment by enabling fine-grained gradient propagation across the entire codebook. Instead of updating only a single selected codeword, FineSID distributes learning signals to all codewords in a soft, differentiable manner. This design promotes globally balanced codebook optimization while strictly preserving semantic consistency, effectively alleviating SID collisions and stabilizing training in large, high-dimensional codebooks. Extensive experiments on multiple public benchmarks demonstrate that FineSID is robust to initialization configurations and consistently improves both codebook utilization and recommendation accuracy. Our work provides a principled, initialization-agnostic solution for semantic identifier learning, advancing the practicality of generative recommendation.

发表机构

  • Tsinghua University(清华大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑