发表机构
Hokkaido University(北海道大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对轮播界面,提出可观测浏览深度的排序老虎机模型,利用浏览深度区分未点击与未展示物品,提出三种算法并证明渐近最优性,实验显示其遗憾低于PBM-UCB。
AI 中文摘要
轮播界面允许推荐系统直接观察用户浏览了多远。该信号区分了已展示但未点击的物品与从未展示的物品,而传统的排序老虎机模型(包括级联模型和基于位置的模型)通常将检查视为潜在变量。我们构建了一个排序老虎机问题,其中学习器呈现一个包含 $L$ 个物品的列表,观察用户的最大浏览深度,并且仅接收该深度之前位置的点击反馈。目标是在未知的物品吸引力向量和浏览深度分布下最大化期望点击次数。我们提出了三种基于 UCB、汤普森采样和 DMED 的算法,所有这些算法仅从观察到的曝光中更新物品统计信息。我们为基于 UCB 的算法推导了实例相关的对数上界,并为基于 DMED 的算法推导了渐近上界,该上界在其参数 $\alpha\downarrow 0$ 时与下界一致,从而在此极限下建立了渐近最优性。在合成的浅层和深层浏览环境中的模拟,以及根据 RecGaze 交互日志参数化的实验表明,OD-TS 达到了与 PBM-TS 相似的最终平均遗憾,而所提出的方法比 PBM-UCB 实现了更低的最终平均遗憾。
英文摘要
Carousel interfaces allow a recommender system to directly observe how far a user has browsed. This signal distinguishes displayed but unclicked items from items that were never displayed, whereas conventional ranking-bandit models, including cascade and position-based models, generally treat examination as latent. We formulate a ranking-bandit problem in which a learner presents a list of $L$ items, observes the user's maximum browsing depth, and receives click feedback only for positions up to that depth. The objective is to maximize the expected number of clicks under an unknown item-attractiveness vector and a browsing-depth distribution. We propose three algorithms based on UCB, Thompson Sampling, and DMED, all of which update item statistics only from observed exposures. We derive an instance-dependent logarithmic upper bound for our UCB-based algorithm and an asymptotic upper bound for our DMED-based algorithm that coincides with the lower bound as its parameter $α\downarrow 0$, establishing asymptotic optimality in this limit. Simulations in synthetic shallow- and deep-browsing environments, together with experiments parameterized from RecGaze interaction logs, show that OD-TS attains final mean regret similar to PBM-TS, while the proposed methods achieve lower final mean regret than PBM-UCB.