arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25104cs.IR

SWIM:生成式重排序中会话监督列表评估的逐步集成度量

SWIM: Step-Wise Integrated Measure for Session-supervised List Evaluation in Generative Re-ranking

Yuanhao Pu, Chenghao Zhang, Chao Feng, Xunyong Yang, Xiang Li, Yongqi Liu, Defu Lian, Kaiqiao Zhan, Kun Gai

首次发表
浏览论文内容

中文总结 AI 辅助

针对工业推荐系统生成式重排序的会话级动态捕捉不足问题,提出SWIM列表级评估器,结合因果掩码Transformer实现高效评估,在重排序任务中显著提升推荐参与度。

中文摘要 AI 辅助

现代工业推荐系统在重排序阶段越来越多地采用生成器-评估器(Generator-Evaluator, G-E)框架。在该范式中,生成器从上游检索和排序模块过滤后的候选池中生成物品列表,而评估器对这些列表打分,为每个请求选择得分最高的列表进行最终曝光。然而在顺序平台(如短视频应用)上,用户会连续消费物品,忽略人工设置的列表边界。传统评估器通过聚合点式值对列表打分,隐含假设曝光相互独立,无法捕捉关键的会话级动态,如上下文依赖、用户连续性以及重复内容的边际效用递减。为弥合这一差距,我们提出SWIM(Step-Wise Integrated Measure,逐步集成度量),这是一种列表级评估器,将用户行为建模为有限时域前缀会话级生存过程。SWIM通过将当前列表对会话级目标的贡献分解为递归生存分布和到达位置条件奖励,估计前缀条件贡献。SWIM利用因果掩码Transformer高效并行估计延续概率和效用,满足严格的工业延迟约束。大量实验表明,SWIM在列表式重排序任务中显著优于基线,在整体推荐参与度上取得大幅提升。

英文摘要

Modern industrial recommender systems have increasingly adopted the Generator-Evaluator (G-E) framework for the re-ranking stage. Within this paradigm, the generator produces candidate item lists from a pool filtered by upstream retrieval and ranking modules, while the evaluator scores these lists and selects the highest-scoring one for final exposure per request. However, on sequential platforms (e.g., short-video apps), users consume items continuously, ignoring artificial list boundaries. Conventional evaluators score lists by aggregating point-wise values, implicitly assuming exposure independence. This fails to capture critical session-level dynamics, such as contextual dependencies, user continuation, and diminishing marginal utility from repetitive content. To bridge this gap, we propose SWIM (Step-Wise Integrated Measure), a list-level evaluator that models user behaviors as a finite-horizon prefix session-level survival process. SWIM estimates the prefix-conditioned contribution of the current list to the session-level objective by factorizing it into a recursive survival distribution and reached-position conditional rewards. Leveraging a causally-masked Transformer, SWIM efficiently estimates continuation probabilities and utilities in parallel, satisfying strict industrial latency constraints. Extensive experiments demonstrate that SWIM significantly outperforms baselines in listwise reranking tasks, yielding substantial improvements in overall recommendation engagement.

补充信息

↑