arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在对数凹随机效用模型中从人类反馈中学习排序

Learning a Ranking from Human Feedback in Log-Concave Random Utility Models

Diego Alovisetti, Marco Mussi, Alberto Maria Metelli

arXiv 2610.07973首次发表:更新:

发表机构

Politecnico di Milano(米兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对对数凹随机效用模型,提出从人类全排序或仅赢家反馈中恢复物品排序的方法,建立样本复杂度下界并开发匹配算法,揭示仅赢家反馈的固有难度。

AI 中文摘要

我们研究根据未知数值效用恢复固定物品集合排序的问题。在与环境的每次交互中,学习者向人类展示物品集合并接收两种类型的比较反馈。在全排序反馈下,每次交互揭示所有物品的带噪声排序,而在仅赢家反馈下,仅揭示排名第一的物品。在两种设置中,我们使用具有对数凹噪声的随机效用模型对人类反馈进行建模,并研究以高概率恢复一个ε-准确排序所需的观测次数。这一新标准仅容忍效用差异小于ε的物品之间的排序错误。对于两种反馈类型,我们建立了最坏情况下的样本复杂度下界,并开发了在对数因子内匹配这些下界的算法。两种算法均不需要噪声分布的知识,仅需要其方差的上界。我们的结果表明,在仅赢家反馈下的排序问题本质上更难,通过揭示样本复杂度对物品集合中最小获胜概率的依赖性来体现。

英文摘要

We study the problem of recovering the ranking of a fixed set of items according to their unknown numerical utilities. At each interaction with the environment, a learner presents the item set to a human and receives comparative feedback of two types. Under full-ranking feedback, each interaction reveals a noisy ranking of all items, whereas under winner-only feedback, it reveals only the item ranked first. In both settings, we model human feedback using a random utility model with log-concave noise and study the number of observations needed to recover an $ε$-accurate ranking with high probability. This novel criterion tolerates ordering errors only between items whose utilities differ by less than $ε$. For both feedback types, we establish worst-case sample-complexity lower bounds and develop algorithms that match these bounds up to logarithmic factors. Neither algorithm requires knowledge of the noise distribution, while only requiring an upper bound on its variance. Our results show that the ranking problem under winner-only feedback is intrinsically harder by exposing the sample complexity dependence on the minimum winning probability across the item set.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑