带Fréchet型尾分布的Follow-the-Perturbed-Leader:对抗式老虎机中的最优性与双世界最优(Best-of-Both-Worlds)特性
Follow-the-Perturbed-Leader with Fréchet-type Tail Distributions: Optimality in Adversarial Bandits and Best-of-Both-Worlds
- Seoul National University(首尔大学)
- Kyoto University(京都大学)
- RIKEN AIP(理化学研究所人工智能项目)
- NEC Corporation(日本电气公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究FTPL策略在K臂老虎机中的最优性,建立了Fréchet型尾分布扰动实现对抗场景O(√(KT))悔值的充分条件,证明其双世界最优能力,为相关猜想和FTRL正则化研究提供新视角。
AI中文摘要:
本文研究了Follow-the-Perturbed-Leader(FTPL,跟随扰动领导者)策略在对抗式和随机K臂老虎机中的最优性。尽管采用多种正则化选择的Follow-the-Regularized-Leader(FTRL,跟随正则化领导者)框架被广泛使用,但依赖随机扰动的FTPL框架尽管本身简洁,却未受到太多关注。在对抗式老虎机中,有猜想认为若扰动服从带Fréchet型尾的分布,FTPL有望达到O(√(KT))的悔值。Honda等人(2023)的近期研究表明,采用形状参数α=2的Fréchet分布的FTPL确实能达到该界,值得注意的是其在随机老虎机中还能获得对数悔值,这意味着FTPL具备双世界最优(Best-of-Both-Worlds, BOBW)能力。然而,该结果仅部分解决了上述猜想,因为他们的分析高度依赖于该形状下Fréchet分布的特定形式。本文中,我们建立了扰动在对抗场景下实现O(√(KT))悔值的充分条件,该条件涵盖Fréchet分布、Pareto分布和Student-t分布等例子。我们还证明了采用特定Fréchet型尾分布的FTPL可以实现BOBW特性。我们的结果不仅有助于通过极值理论视角解决现有猜想,还可能通过FTPL到FTRL的映射,为理解FTRL中正则化函数的作用提供洞见。
英文摘要:
This paper studies the optimality of the Follow-the-Perturbed-Leader (FTPL) policy in both adversarial and stochastic $K$-armed bandits. Despite the widespread use of the Follow-the-Regularized-Leader (FTRL) framework with various choices of regularization, the FTPL framework, which relies on random perturbations, has not received much attention, despite its inherent simplicity. In adversarial bandits, there has been conjecture that FTPL could potentially achieve $\mathcal{O}(\sqrt{KT})$ regrets if perturbations follow a distribution with a Fréchet-type tail. Recent work by Honda et al. (2023) showed that FTPL with Fréchet distribution with shape $α=2$ indeed attains this bound and, notably logarithmic regret in stochastic bandits, meaning the Best-of-Both-Worlds (BOBW) capability of FTPL. However, this result only partly resolves the above conjecture because their analysis heavily relies on the specific form of the Fréchet distribution with this shape. In this paper, we establish a sufficient condition for perturbations to achieve $\mathcal{O}(\sqrt{KT})$ regrets in the adversarial setting, which covers, e.g., Fréchet, Pareto, and Student-$t$ distributions. We also demonstrate the BOBW achievability of FTPL with certain Fréchet-type tail distributions. Our results contribute not only to resolving existing conjectures through the lens of extreme value theory but also potentially offer insights into the effect of the regularization functions in FTRL through the mapping from FTPL to FTRL.