AI 中文总结
本博士研究针对轮播推荐中排序与呈现耦合带来的评估挑战,开发点击模型框架、轮播特定离线指标等,旨在建立关联轮播点击建模与离线离策略评估的连贯体系以可靠改进推荐策略。
AI 中文摘要
轮播界面被广泛应用于现代推荐系统。与呈现单个排序列表的传统界面不同,轮播界面同时向用户呈现多个排序列表,作为相互堆叠的可水平滑动行。在这种设计中,排序与二维布局紧密关联。因此,用户行为不仅受物品偏好影响,还受行组织、视口约束和物品上下文的塑造。排序与呈现之间的这种紧密耦合使用户反馈的解释变得复杂,为推荐评估带来了新挑战。我的博士研究旨在通过重新思考轮播点击的建模方式,以及如何从记录的交互数据中评估轮播推荐策略来应对这些挑战。迄今为止,我已研究了用户与轮播界面的交互方式,并开发了一种点击模型设计框架,该框架优先考虑观测变量之间的数学关系而非潜在行为假设。基于这些成果,我正在进行的工作包括一个使用离散选择模型将点击表示为选择的项目,以及一个开发轮播特定离线指标的项目。下一步,我计划开发离策略评估方法,以从记录的交互中估计推荐策略的性能。总体而言,本论文的预期贡献是形成一套连贯的研究体系,将轮播点击建模与离线和离策略评估关联起来,从而能够更可靠地改进轮播推荐策略。
英文摘要
Carousel interfaces are widely used in modern recommendation systems. Unlike traditional interfaces that present a single ranked list, carousels simultaneously present several ranked lists to the user, as horizontally swipeable rows stacked on top of each other. In this design, the rankings are closely tied to the two-dimensional layout. Consequently, user behavior is shaped not only by item preference, but also by row organization, viewport constraints, and item context. This tight coupling between ranking and presentation complicates the interpretation of user feedback, introducing new challenges for recommendation evaluation. My PhD research aims to address these challenges by rethinking how carousel clicks are modeled and how carousel recommendation policies can be evaluated from logged interaction data. So far, I have studied how users interact with carousel interfaces and developed a click model design framework that prioritizes mathematical relationships between observed variables over latent behavioral assumptions. Building on these results, my ongoing work includes a project using discrete choice models to represent clicks as choices, alongside a project that develops carousel-specific offline metrics. As a next step, I plan to develop off-policy evaluation methods that estimate the performance of recommendation policies from logged interactions. Taken together, the expected contribution of my thesis is a connected body of work that links carousel click modeling with offline and off-policy evaluation, so that carousel recommendation policies can be improved more reliably.