arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.16382stat.MLcs.LG

稀疏非参数上下文老虎机

Sparse Nonparametric Contextual Bandits

  • Universitat Pompeu Fabra(庞培法布拉大学)
  • Delft University of Technology(代尔夫特理工大学)
  • Bocconi University(博科尼大学)

机构由 AI 辅助整理,请以论文原文为准。

Hamish Flynn, Julia Olkhovskaya, Paul Rognon-Vael

更新

AI总结:

该研究探讨了非参数上下文老虎机中稀疏性对最小最大遗憾的影响,提出了一种新的减少方法并展示了Feel-Good Thompson采样算法的改进性能。

AI中文摘要:

我们研究了非参数上下文老虎机问题中稀疏性的好处,其中候选特征集是可数或不可数无限的。我们的贡献是双方面的。首先,通过一种新颖的将问题转化为序列多臂老虎机问题的减少方法,我们提供了最小最大遗憾的下界,这表明在该设置中多项式依赖于动作数量通常是不可避免的。其次,我们证明了一种Feel-Good Thompson采样算法的变种在遗憾界上达到了我们的下界,仅在时间跨度上存在对数因子,并且在候选特征的有效数量上具有对数依赖性。当我们将我们的结果应用于核化和神经上下文老虎机时,我们发现稀疏性在时间跨度足够大相对稀疏性和动作数量时能够带来更好的遗憾界。

英文摘要:

We study the benefits of sparsity in nonparametric contextual bandit problems, in which the set of candidate features is countably or uncountably infinite. Our contribution is two-fold. First, using a novel reduction to sequences of multi-armed bandit problems, we provide lower bounds on the minimax regret, which show that polynomial dependence on the number of actions is generally unavoidable in this setting. Second, we show that a variant of the Feel-Good Thompson Sampling algorithm enjoys regret bounds that match our lower bounds up to logarithmic factors of the horizon, and have logarithmic dependence on the effective number of candidate features. When we apply our results to kernelised and neural contextual bandits, we find that sparsity enables better regret bounds whenever the horizon is large enough relative to the sparsity and the number of actions.

补充信息

↑