arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

带共形动作集的强化学习:在序列推荐中的应用

Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation

Wenwen Si, Honghao Wei

arXiv 2610.08743首次发表:更新:

发表机构

University of Pennsylvania; Washington State University(宾夕法尼亚大学; 华盛顿州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对序列推荐中固定候选集大小的问题,提出带校准剪枝的强化学习(RLCP),通过评论家分数和在线阈值自适应调整动作集,在19种配置中至少一种变体实现最高目录多样性,达最强基线的1.11至5.21倍。

AI 中文摘要

序列推荐器通常使用固定的候选集大小,即使会话中有用备选方案的数量会发生变化。我们提出了带校准剪枝的强化学习(RLCP),该方法利用评论家分数和在线阈值自适应地调整保留的动作集。该阈值根据二元反馈进行更新,该反馈指示该集合是否包含代理目标中的某个动作。我们证明了在自适应轨迹上观察到的代理未命中率的确定性界。为了量化剪枝对奖励的影响,我们将价值损失精确分解为过滤损失和选择损失。在显式代理和评论家近似条件下,该分解产生一个有限的会话奖励界,该界同时考虑了不完美的选择和集合截断,而无需学习参数收敛。在KuaiRand-Pure和MovieLens 1M上的实验将两种RLCP实现与四种强化学习基线进行了比较。在19种配置中的每一种中,至少一种RLCP变体实现了最高的目录多样性,达到最强基线的1.11倍至5.21倍,同时具有有竞争力的会话深度且保留集不更大。

英文摘要

Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching $1.11\times$ to $5.21\times$ that of the strongest baseline, with competitive session depth and no larger retained sets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑