AI 中文总结
针对生成式推荐器中物品分词引发的物品内部注意力过载问题,提出SST方法,通过IST与BCA优化表示,在多数据集与骨干模型上实现性能提升。
AI 中文摘要
在生成式推荐系统中,物品通常被分词为固定长度的语义ID序列,用于自回归式的下一个物品预测。但在用户-上下文建模中,这种细粒度表示会引发物品内部注意力过载:过多注意力被用于低层次的物品内部依赖,而非高层次的物品间行为转换。为解决该问题,我们提出语义子词分词(Semantic Subword Tokenization, SST),将历史物品表示为可变长度的语义子词,同时保持目标解码的固定长度。SST首先应用物品级子词分词(Item-level Subword Tokenization, IST),将稳定的相邻原子令牌合并为紧凑的语义子词令牌,从而减少编码器中的物品内部重组;随后引入行为诱导共现增强(Behavior-induced Co-occurrence Augmentation, BCA),注入粗粒度的语义前缀转换信号,引导释放的建模能力转向物品间行为规律。在三个公开数据集和三个生成式推荐器骨干模型上进行的大量实验表明,SST相较于固定长度和可迁移可变长度SID基线取得了经验性提升。代码可在该httpsURL获取。
英文摘要
In generative recommender systems, items are typically tokenized into fixed-length semantic ID sequences for autoregressive next-item prediction. However, for user-context modeling, this fine-grained representation triggers Intra-item Attention Overload: excessive attention is spent on low-level intra-item dependencies rather than high-level inter-item behavioral transitions. To address this, we propose Semantic Subword Tokenization (SST), which represents historical items as variable-length semantic subwords while preserving fixed-length target decoding. SST first applies Item-level Subword Tokenization (IST) to merge stable adjacent atom tokens into compact semantic subword tokens, thereby reducing intra-item reassembly in the encoder. It then introduces Behavior-induced Co-occurrence Augmentation (BCA) to inject coarse-grained semantic prefix transition signals, guiding the freed modeling capacity toward inter-item behavioral regularities. Extensive experiments on three public datasets and three generative recommender backbones show empirical improvements of SST over fixed-length and transferable variable-length SID baselines. Code is available at https://github.com/mxrcandy/Semantic-Subword-Tokenization.
Comments13 pages, 6 figures, 8 tables. Accepted to CIKM 2026