发表机构
College of Computer Science, Nankai University; College of Cryptology and Cyber Science, Nankai University(南开大学计算机学院; 南开大学密码与网络科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于SID的生成式推荐中存在的令牌频率偏差问题,提出FSGR框架,通过优化SID构建与训练策略缓解偏差,在保持准确率的同时提升了约20%的基尼公平性。
AI 中文摘要
基于语义ID(SID)的生成式推荐近期取得了显著成功,但现有方法存在一个此前被忽视的公平性问题,我们将其命名为“令牌频率偏差”,即高频SID令牌被系统性地过度预测,而低频SID令牌则被预测不足。该偏差源于SID构建过程中语义码本的不平衡,以及推荐训练过程中的流行度偏差与最大似然估计目标的共同作用,最终导致不同物品类别间的曝光不公平。现有SID方法主要聚焦于提升码本质量,却忽视了令牌频率不平衡对下游推荐公平性的影响;而大语言模型(LLM)去偏方法因SID令牌具有层级语义,直接应用于基于SID的推荐时往往效果不佳。为解决该问题,我们提出了FSGR,一种针对基于SID的生成式推荐的公平性优化框架。在SID构建阶段,FSGR采用基于最优传输(OT)的分配优化与双标准重锚机制,以形成更平衡的SID表示空间;在推荐训练阶段,它采用两阶段训练策略,并引入层级频率校准以进行特定层的公平性微调。在三个公共数据集上针对三个骨干模型开展的实验表明,FSGR可缓解令牌频率偏差,在保持推荐准确率竞争力的同时,实现了超过20%的平均基尼公平性提升。
英文摘要
Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previously overlooked fairness issue, which we term \textbf{Token Frequency Bias}, where high-frequency SID tokens are systematically over-predicted while low-frequency SID tokens are under-predicted. This bias originates from the combined effects of imbalanced semantic codebooks during SID construction, and popularity bias together with the maximum likelihood estimation objective during recommendation training, resulting in unfair exposure across item categories. Existing SID methods mainly focus on improving codebook quality and overlook the impact of token frequency imbalance on downstream recommendation fairness, while LLM debiasing methods often yield suboptimal results when directly applied to SID-based recommendation, due to the hierarchical semantics of SID tokens. To address this issue, we propose \textbf{FSGR}, a fairness optimization framework for SID-based generative recommendation. During SID construction, FSGR employs OT-based Assignment Optimization and Dual-Criteria Re-anchor mechanism to form a more balanced SID representation space. During recommendation training, it adopts a two-stage training strategy and introduces Hierarchical Frequency Calibration for layer-specific fairness fine-tuning. Experiments on three public datasets with three backbone models demonstrate that FSGR mitigates token frequency bias and delivers an average Gini fairness improvement of over 20\% while maintaining competitive recommendation accuracy.