基于曝光的排序强化学习
Exposure-Based Reinforcement Learning to Rank
浏览论文内容
中文总结 AI 辅助
研究排序学习的强化学习方法,针对标准RL在LTR中低效及现有方法复杂的问题,通过聚焦方差减少和GPU计算,提出新的梯度估计抽象,实现与自动微分无缝集成,实验证明新方法收敛快、性能高且无额外计算成本。
中文摘要 AI 辅助
排序学习(LTR)的强化学习(RL)方法可优化任何排序目标。但标准RL在LTR设置中因动作空间巨大而低效且计算成本高。现有方法通过自定义梯度计算算法提高效率,但实现复杂且与自动微分冲突。本文重新考虑LTR的RL,避免依赖自定义梯度,聚焦方差减少和GPU计算。通过基线校正和部分边缘化实现高样本效率,提出文档曝光分布后的梯度估计抽象,可与自动微分无缝集成。实验结果表明新方法收敛更快、排序性能更高,无额外计算成本,而现有方法存在稳定性问题。
英文摘要
Reinforcement learning (RL) methods for learning-to-rank (LTR) can optimize (almost) any ranking goal, e.g., from precision or discounted cumulative gain to fairness-of-exposure or ranking distillation. However, standard RL is ineffective and computationally costly due to the enormous action space in LTR settings. Existing methods reach computational efficiency through custom gradient computation algorithms, but they are very complex to implement and often clash with auto-differentiation. Consequently, existing RL for LTR is not attractive to many practitioners. We reconsider RL for LTR while actively avoiding reliance on custom gradients. Contrary to the existing approaches, we focus on variance reduction and GPU computation. In doing so, we discover that high sample-efficiency can be reached through baseline corrections and partial marginalization. Furthermore, we propose an abstraction that places gradient estimation behind a document-exposure distribution, this enables seamless plug-and-play integration with auto-differentiation. Thereby, one only has to implement a loss as a differentiable function of exposure and RL for LTR can optimize it using auto-differentiation. Our experimental results reveal that our new exposure-based RL for LTR approach converges considerably faster and at significantly higher ranking performance than existing custom gradients, with no additional costs in computation time when using GPUs. In contrast, existing custom gradients result in severe stability issues when converging over many epochs, which never occur for our methods. Thus, we considerably improve RL for LTR methodology by increasing its effectiveness, efficiency, and ease of application.
发表机构
- University of Amsterdam(阿姆斯特丹大学)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。