发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对偏好异质性场景,探究成对比较排序所需的重复次数,提出三种算法将所需次数从Ω(1/Δ²)降至最优的O(log(1/Δ)),其中俄罗斯轮盘算法期望仅需O(1)次,实验验证了算法性能。
AI 中文摘要
我们研究当偏好在用户和任务间存在差异时,基于成对比较的总体平均效用排序模型。现有研究表明,即使用户数量任意多,每个用户仅一次比较也不足以识别具有最高平均效用的备选方案(Golz等人,2025)。我们探究在每个用户-任务情境下,需要多少重复比较才能实现排序恢复。在具有固定逆温度的异质性布拉德利-特里(Bradley-Terry)模型下,我们首先提出一种基于朴素最大似然估计(MLE)的算法,该算法需要每个情境下Ω(1/Δ²)次重复比较以确保排序恢复。随后,我们提出两种基于MLE的变体算法,以及一种随机俄罗斯轮盘(Russian Roulette)式算法,这三种算法可在每个情境下使用O(log(1/Δ))次重复比较实现排序恢复,且我们证明这种对数依赖是最优的。尽管存在最坏情况要求,我们的俄罗斯轮盘算法在每个情境下仅需O(1)次期望比较。合成实验与基于Arena数据的半合成实验,在不同偏好异质性水平和不同情境分布的设置下,对这四种算法进行了比较。
英文摘要
We study ranking models by population-average utility from pairwise comparisons when preferences vary across users and tasks. Prior work shows that a single comparison per user can be insufficient to identify the alternative with the highest average utility, even with arbitrarily many users (Golz et al., 2025). We investigate how many repeated comparisons within each user-task context are necessary and sufficient for ranking recovery. Under a heterogeneous Bradley-Terry model with fixed inverse temperature, we start with a naive MLE-based algorithm that requires $Ω(1/Δ^2)$ repeated comparisons per context to ensure ranking recovery. We then present two MLE-based variants and a randomized Russian Roulette-style algorithm that recover the ranking using $O(\log(1/Δ))$ repeated comparisons per context, and we prove that this logarithmic dependence is optimal. Despite this worst-case requirement, our Russian Roulette algorithm uses only $O(1)$ comparisons per context in expectation. Synthetic experiments and semi-synthetic experiments based on Arena data compare the four algorithms in settings with varying levels of preference heterogeneity and under varying context distributions.
Comments53 pages, 4 figures