arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25268cs.IRcs.AI

用于排序的结构感知相对策略优化

Structure-aware Relative Policy Optimization for Ranking

Yiteng Tu, Weihang Su, Zitao Su, Yiqun Liu, Min Zhang, Qingyao Ai

首次发表
浏览论文内容

中文总结 AI 辅助

研究排序问题,提出结构感知相对策略优化框架SRPO,通过顶部加权肯德尔-陶距离测量排列差异并归一化奖励差异,强调高效局部优化,实验证明该方法能提高列表排序的有效性和稳定性。

中文摘要 AI 辅助

排序是现代信息访问系统的基本组成部分。强化学习为直接优化粗粒度反馈和基于完整排序列表定义的系统级目标提供了一个灵活框架。但现有基于强化学习的排序方法通常将每个采样排列视为原子输出,主要通过标量奖励评估,忽视了不同排序列表间的结构关系。为此提出SRPO,一种用于列表排序的结构感知相对策略优化框架。SRPO使用顶部加权肯德尔-陶距离测量采样排列之间的差异,并通过相应距离对其成对奖励差异进行归一化。实验结果表明,明确建模排列级差异可提高列表排序的有效性和稳定性。

英文摘要

Ranking is a fundamental component of modern information access systems. Reinforcement learning (RL) provides a flexible framework for directly optimizing coarse-grained feedback and system-level objectives defined over the complete ranking list. However, existing RL-based ranking methods typically treat each sampled permutation as an atomic output and evaluate it primarily through a scalar reward, overlooking the structural relationships among different ranking lists. Consequently, permutations with similar rewards but substantially different permutation patterns may receive comparable optimization signals, potentially leading to inaccurate credit assignment and overly aggressive policy updates. To address this limitation, we propose SRPO, a \textbf{S}tructure-aware \textbf{R}elative \textbf{P}olicy \textbf{O}ptimization framework for listwise ranking. SRPO measures the discrepancy between sampled permutations using a top-weighted Kendall-tau distance and normalizes their pairwise reward differences by the corresponding distances. It quantifies the reward improvement per unit of ranking change, thereby emphasizing efficient local refinements, particularly those involving top-ranked positions. Experimental results across two ranking scenarios demonstrate that explicitly modeling permutation-level differences improves the effectiveness and stability of listwise ranking, with particularly favorable performance in limited-feedback and complex list-level optimization settings.

发表机构

  • Tsinghua University(清华大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

↑