饱和不敏感的决斗式赌博机与通用函数逼近
Saturation-Insensitive Dueling Bandits with General Function Approximation
AI总结:
针对决斗式赌博机在偏好模型饱和时的样本复杂度问题,提出SI-CDB算法,通过启发式对手选择消除逆导数因子,实现饱和不敏感学习,并借助局部Eluder维度框架获得近最优理论保证。
AI中文摘要:
我们研究了在Bradley-Terry-Luce(BTL)偏好模型下,具有通用函数逼近的上下文决斗式赌博机问题。该设置中的一个关键挑战是偏好模型的饱和性:当当前奖励模型已经能够以高置信度区分两个动作时,产生的偏好反馈变得信息量较弱,使得进一步改进奖励估计变得困难。因此,现有的样本复杂度分析通常依赖于逆导数因子 $1 / \sigma'[\Delta_{r^\ast}]$,当链接函数 $\sigma$ 对于大的奖励差距 $\Delta_{r^\ast}$ 饱和时,该因子可能过大。为了解决这个问题,我们引入了 `SI-CDB`,一种使用精心设计的启发式方法选择对手臂的算法。这种设计实现了饱和不敏感的奖励学习,并为线性奖励类别恢复了接近最优的依赖性,消除了不利的 $1/\sigma'(\cdot)$ 因子。我们分析的核心是一个针对具有通用函数逼近的决斗式赌博机量身定制的局部Eluder维度框架。我们的理论结果还解释了为什么双臂遗憾分析对于提高决斗式赌博机中单臂性能至关重要。
英文摘要:
We study contextual dueling bandits with general function approximation under the Bradley-Terry-Luce (BTL) preference model. A key challenge in this setting is the saturation of the preference model: when the current reward model can already distinguish two actions with high confidence, the resulting preference feedback becomes weakly informative, making it difficult to further improve reward estimation. Consequently, existing sample-complexity analyses often depend on the inverse-derivative factor $1 / σ'[Δ_{r^\ast}]$ which can be prohibitively large when the link function $σ$ saturates for large reward gaps $Δ_{r^\ast}$. To address this issue, we introduce `SI-CDB`, an algorithm that selects opponent arms using a carefully designed heuristic for arm selection. This design enables saturation-insensitive reward learning and recovers the near-optimal dependence for linear reward classes, eliminating the unfavorable $1/σ'(\cdot)$ factor. The core of our analysis is a localized Eluder dimension framework tailored to dueling bandits with general function approximation. Our theoretical results also explain why two-arm regret analysis is crucial for improving single-arm performance in dueling bandits.