解耦赌博机的跟随扰动领导者:两全其美与实用性
Follow-the-Perturbed-Leader for Decoupled Bandits: Best-of-Both-Worlds and Practicality
- Seoul National University, Seoul, Korea(首尔国立大学,韩国首尔)
- Korea Institute of Science and Technology, Seoul, Korea(韩国科学技术院,韩国首尔)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对解耦多臂赌博机问题,提出一种高效的跟随扰动领导者策略,在随机环境下实现常数遗憾,在对抗环境下实现最优O(√KT)遗憾,且避免了凸优化和重采样过程,显著降低计算成本。
AI中文摘要:
我们研究了解耦多臂赌博机问题,其中学习者在每一轮分别选择一个臂进行探索,并选择另一个可能不同的臂进行利用。在此设置中,探索臂的损失被观察到但不承担,而利用臂的损失被承担但不被观察到。我们提出了一种高效的跟随扰动领导者(FTPL)策略,该策略在随机环境下实现常数遗憾,在对抗环境下实现最优$O(\sqrt{KT})$遗憾,从而获得两全其美(BOBW)保证。我们方法的一个关键特征是它完全避免了先前BOBW策略所需的凸优化以及FTPL赌博机策略中通常使用的重采样过程。这使得FTPL能够充分发挥其计算效率优势,大幅降低计算成本。我们通过实验证实,我们的策略不仅提高了运行时间,而且在两种环境下都表现出优越的遗憾性能。
英文摘要:
We study the decoupled multi-armed bandit problem, where the learner separately selects one arm for exploration and one, possibly different, arm for exploitation at each round. In this setting, the loss of the explored arm is observed but not incurred, whereas the loss of the exploited arm is incurred without being observed. We propose an efficient Follow-the-Perturbed-Leader (FTPL) policy that achieves Best-of-Both-Worlds (BOBW) guarantee with constant regret in the stochastic regime and optimal $O(\sqrt{KT})$ regret in the adversarial regime. A key feature of our method is that it completely avoids both the convex optimization required by prior BOBW policies and the resampling procedures typically used in FTPL bandit policies. This allows FTPL to fully realize its computational efficiency advantages, leading to substantial reductions in computational cost. We empirically confirm that our policy not only improves the runtime but also demonstrates superior regret performance in both regimes.