从成对比较中进行最优的前\(k\)项识别
Optimal Top-$k$ Identification from Pairwise Comparisons
AI总结:
研究从成对比较中进行固定置信度前\(k\)项识别的主动学习问题,刻画下界结构为鞍点问题并开发渐近最优算法,构造自适应比较分配算法并证明其渐近最优,最小化比较预期数量。
AI中文摘要:
我们研究了从有噪声的成对比较中进行固定置信度的前\(k\)项识别的主动学习问题。在此问题中,算法依次选择项目对进行比较,观察结果,并在能够以至多\(\delta\)的错误概率返回前\(k\)项集合时停止。目标是设计这样一个\(\delta\)正确的过程,使其最小化比较的预期数量(样本复杂度)。该问题属于更广泛的关于带反馈模型中固定置信度纯探索的文献,常见目标是渐近最优性:当\(\delta \to 0\)时,算法的预期样本复杂度与信息论下界匹配。针对一系列固定置信度纯探索问题已开发出渐近最优过程,但据我们所知,对于前\(1\)项,或更一般地从潜在效用模型下的成对比较中进行前\(k\)项识别,尚未建立渐近最优算法。在此设置下,我们开发了这样一种算法。我们刻画了下界的结构并将其表述为鞍点问题。这种结构使得能够通过计算效率高的原始对偶过程在线学习渐近最优的比较分配。然后我们构造了一种自适应比较分配算法,跟踪由原始对偶过程学习到的分配,并证明它是渐近最优的。
英文摘要:
We study the active learning problem of fixed-confidence top-$k$ identification from noisy pairwise comparisons. In this problem, an algorithm sequentially chooses pairs of items to compare, observes the outcomes, and stops when it can return the set of top-$k$ items with error probability at most $δ$. The objective is to design such a $δ$-correct procedure that minimizes the expected number of comparisons (the sample complexity). This problem falls within the broader literature on fixed-confidence pure exploration in bandit models, where a common target is asymptotic optimality: the algorithm's expected sample complexity matches the information theoretic lower bound as $δ\to 0$. Asymptotically optimal procedures have been developed for a range of fixed-confidence pure-exploration problems, however to the best of our knowledge, for top-$1$, or more generally top-$k$ identification from pairwise comparisons under latent utility models an asymptotically optimal algorithm has not been established. In this setting, we develop such an algorithm. We characterize the structure of the lower bound and formulate it as a saddle-point problem. This structure enables a computationally efficient primal-dual procedure that learns the asymptotically optimal comparison allocation online. We then construct an adaptive comparison-allocation algorithm that tracks the allocation learned by the primal-dual procedure and prove it is asymptotically optimal.