关于基于偏好的老虎机问题的复杂性
On the Complexity of Preference-Based Bandits
浏览论文内容
中文总结 AI 辅助
本文研究基于偏好的老虎机问题,提出局部敏感eluder维度作为复杂度度量,并设计GINOP算法,实现与直接奖励学习同等统计效率的遗憾界。
中文摘要 AI 辅助
我们研究具有一般奖励函数类别的基于偏好的老虎机问题,其中学习者顺序选择一对臂并观察由Bradley-Terry模型控制的二元偏好反馈。这种设置自然出现在推荐系统、锦标赛排名和从人类反馈中学习等应用中,在这些应用中,相对偏好比绝对奖励更容易获得。观察模型继承了逻辑老虎机处理问题依赖常数κ的挑战,该常数解释了链接函数的非线性,并且可以任意大。此外,先前的工作主要集中在线性或核化奖励模型上,排除了使用更丰富的函数类别。为了解决这些限制,我们考虑一般奖励函数类别,并引入了局部敏感eluder维度,这是一种针对偏好反馈的逻辑结构量身定制的新颖复杂度度量,可以在不依赖κ的不利影响的情况下产生细粒度的遗憾界。基于这一概念,我们提出了GINOP(通用信息性乐观)算法,该算法构建对数损失置信集并联合选择臂对以平衡乐观和信息性探索。我们建立了一阶遗憾界,与先前结果所暗示的相反,这表明使用偏好反馈进行学习在统计效率上与直接奖励观察学习一样高效。最后,我们通过与竞争基线的实证评估来证实我们的理论发现。
英文摘要
We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandit challenge of handling the problem-dependent constant $κ$, which accounts for the non-linearity of the link function and can grow arbitrarily large. Moreover, prior work has predominantly focused on linear or kernelized reward models, precluding the use of richer function classes. To address these limitations, we consider general reward function classes and introduce the \emph{locally sensitive eluder dimension}, a novel complexity measure tailored to the logistic structure of preference feedback that yields fine-grained regret guarantees without unfavorable dependence on $κ$. Building on this notion, we propose \textbf{GINOP} (Generic INformative OPtimism), an algorithm that constructs log-loss confidence sets and jointly selects arm pairs to balance optimism and informative exploration. We establish a first-order regret bound that, in contrast with what previous results suggest, demonstrates that learning with preference feedback is as statistically efficient as learning from direct reward observation. Finally, we corroborate our theoretical findings with empirical evaluations against competitive baselines.
发表机构
- Criteo AI Lab(Criteo人工智能实验室)
- CREST(经济与统计研究中心)
- ENSAE(法国国立统计与经济管理学校)
机构由 AI 辅助整理,请以论文原文为准。