arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.05098stat.MLcs.LG

面向范数未知老虎机的经验界信息导向采样

Empirical Bound Information-Directed Sampling for Norm-Agnostic Bandits

  • Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

Piotr M. Suder, Eric Laber

更新

AI总结:

本文提出一种无需事先知道参数范数上界的频率学派信息导向采样算法,通过迭代细化高概率范数界并混合信息增益准则,在线性异方差次高斯老虎机中实现优于现有 IDS 和 UCB 的遗憾性能。

AI中文摘要:

信息导向采样(IDS)是解决老虎机问题的一个强大框架,在贝叶斯和频率学派设定下都展现出优异的结果。然而,频率学派 IDS 与许多其他老虎机算法一样,要求事先知道控制奖励模型的真实参数向量范数的一个(相对)紧的上界,才能取得良好性能。不幸的是,这一要求在实践中很少得到满足。正如我们所展示的,使用校准不当的上界会导致显著的遗憾累积。为了解决这一问题,我们提出了一种新颖的频率学派 IDS 算法,该算法利用不断积累的数据迭代地细化真实参数范数的高概率上界。我们专注于具有异方差次高斯噪声的线性老虎机设定。我们的方法利用相关信息增益准则的混合,来平衡旨在收紧估计参数范数界的探索与直接搜索最优动作。我们为算法建立了不依赖于初始假设参数范数界的遗憾界,并证明我们的方法优于最先进的 IDS 和 UCB 算法。

英文摘要:

Information-directed sampling (IDS) is a powerful framework for solving bandit problems which has shown strong results in both Bayesian and frequentist settings. However, frequentist IDS, like many other bandit algorithms, requires that one have prior knowledge of a (relatively) tight upper bound on the norm of the true parameter vector governing the reward model in order to achieve good performance. Unfortunately, this requirement is rarely satisfied in practice. As we demonstrate, using a poorly calibrated bound can lead to significant regret accumulation. To address this issue, we introduce a novel frequentist IDS algorithm that iteratively refines a high-probability upper bound on the true parameter norm using accumulating data. We focus on the linear bandit setting with heteroskedastic subgaussian noise. Our method leverages a mixture of relevant information gain criteria to balance exploration aimed at tightening the estimated parameter norm bound and directly searching for the optimal action. We establish regret bounds for our algorithm that do not depend on an initially assumed parameter norm bound and demonstrate that our method outperforms state-of-the-art IDS and UCB algorithms.

↑