多臂伯努利老虎机的最小最大单臂停止策略
Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping
浏览论文内容
中文总结 AI 辅助
本文提出基于单臂停止问题最小最大解的索引策略,用于有限时域伯努利多臂老虎机,实现分布无关的遗憾界,并在实验中优于基准策略。
中文摘要 AI 辅助
我们针对有限时域伯努利多臂老虎机问题,基于单臂老虎机(SAB)问题的最小最大解,开发了一种索引策略。每个SAB问题涉及在未知的伯努利臂和已知奖励之间进行选择。我们证明,在所有非预期策略上最小化SAB问题的最坏情况遗憾,可精确地表述为一个半无限线性规划。由此产生的停止策略提供了一种自然的臂比较方式:策略继续采样的已知奖励越高,未知臂就越有前景。我们将这一直觉转化为基于累积继续概率的索引,并带有单调调整和奖励缺口上限。通过将索引误差与单臂停止策略的遗憾联系起来,我们为K个臂和时域T建立了一个无分布遗憾界:$4.45\sqrt{KT}+10.75K$。该界与文献中建立的最小最大最优遗憾阶相匹配。该保证通过伯努利随机化扩展到支持在$[0,1]$上的奖励。我们还提供了一种具有量化近似损失的有限网格实现。在数值实验中,基于SAB的索引策略在所有评估的臂数和时域上,其最坏情况遗憾均低于所有测试的基准策略,同时在双臂设置中与基于网格的MAB最小最大策略紧密匹配。
英文摘要
We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems. Each SAB problem involves choosing between an unknown Bernoulli arm and a known reward. We show that minimizing worst-case regret of SAB problems over all non-anticipative policies admits an exact semi-infinite linear programming formulation. The resulting stopping policies offer a natural way to compare arms: the higher the known reward against which a policy continues sampling, the more promising the unknown arm. We turn this intuition into indices based on cumulative continuation probabilities, with a monotone adjustment and a reward-shortfall cap. By relating index errors to the regret of single-arm stopping policies, we establish a distribution-free regret bound of $4.45\sqrt{KT}+10.75K$ for $K$ arms and horizon $T$. This bound matches the minimax-optimal regret order established in the literature. The guarantee extends to rewards supported on $[0,1]$ through Bernoulli randomization. We also provide a finite-grid implementation with quantified approximation loss. In numerical experiments, the SAB-based index policy achieves lower worst-case regret than every tested benchmark policy across all evaluated numbers of arms and horizons, while closely matching the grid-based MAB minimax policy in the two-arm setting.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- The University of Sydney(悉尼大学)
- École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)
- Imperial College Business School(帝国理工学院商学院)
机构由 AI 辅助整理,请以论文原文为准。