AI 中文总结
针对自回归生成式推荐重排序的局限,提出对空间生成(PSG)方法,理论上可提升生成速度与次优性,在快手部署后实现每用户停留时间0.178%的提升。
AI 中文摘要
现代推荐系统采用生成器-评估器(G-E)框架进行列表级重排序:生成器从候选项中生成序列,评估器对序列进行级评分以筛选出最优序列用于曝光。自回归(AR)作为生成式推荐的骨干存在两个局限:其一,其复杂度随列表长度线性增长,迫使系统在严格延迟约束下生成更少列表,从而限制了探索性;其二,教师强制训练会导致训练-测试不匹配,累积误差随长度增加而恶化,降低了生成质量。为解决这些问题,我们提出对空间生成(PSG),这是一种将生成原子从单个项目提升为有序项目对的重构方法。给定n个候选项,PSG在每个请求上操作于大小为n(n-1)的对词汇表,仅生成L/2个token。对token表示由预训练的对token表示模块实时生成,该模块在大规模曝光日志上优化,消除了原本会困扰二次规模词汇表的数据稀疏性。我们建立了三个理论保证:(i)PSG与项目空间生成是双射的,诱导出等价的序列分布族,因此不损失表达能力;(ii)在中等设置下,对token空间的生成理论上实现约2至4倍的加速,在实际工业环境设置中为1.83倍;(iii)在仅结果奖励下,PSG的最坏情况次优性被O((L/2)²ε̄)界定,相比项目空间生成实现了近4倍的提升。除基于基准的验证外,PSG已在快手部署,使平台的每用户停留时间提升0.178%,该平台拥有超过4亿日活跃用户。
英文摘要
Modern recommender systems adopt Generator-Evaluator (G-E) for list-wise reranking: a generator produces sequences from candidates and an evaluator scores them at sequence-level to filter out the optimal one for exposure. Auto-Regressive(AR), working as the backbone for generative recommendation, suffers two limitations. First, its complexity grows linearly with list length, forcing the system to generate fewer lists under rigorous latency constraints and thus limiting exploration. Second, teacher-forcing creates a train-test mismatch; cumulative errors worsen with length and degrade quality. To address these problems, we propose Pair-Space Generation (PSG), a reformulation that elevates the generation atom from individual items to ordered item pairs. Given $n$ candidate items, PSG operates over pair vocabulary of size $n(n-1)$ per request, generates only $L/2$ tokens. Pair token representations are produced on-the-fly by a pretrained pair-token representation module optimized over large scale exposure logs, eliminating the data sparsity that would otherwise plague a quadratic sized vocabulary. We establish three theoretical guarantees: (i) PSG is bijective with item-space generation and induces an equivalent family of sequence distributions, thus incurring no loss of expressiveness; (ii) generation in pair-token space achieves approximately a $2\times$ to $4\times$ speedup theoretically under moderate settings and $1.83\times$ in the real industrial environmental settings; and (iii) under outcome-only rewards, the worst-case suboptimality of PSG is bounded by $O((L/2)^2 \barε)$, representing a nearly $4\times$ improvement over item-space generation. Beyond benchmark-based validation, PSG has also been deployed on Kuaishou, delivering a 0.178\% lift in per-user stay time on the platform, which serves over 400 million daily active users.
Comments13 pages, 3 figures