arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.12490cs.LGstat.ML

关于组合半老虎机问题中扰动跟随领导者策略的注记

Note on Follow-the-Perturbed-Leader in Combinatorial Semi-Bandit Problems

  • Kyoto University(京都大学)
  • RIKEN AIP(理化学研究所(Advanced Institute for Scientific Technology))

机构由 AI 辅助整理,请以论文原文为准。

Botao Chen, Junya Honda

更新

AI总结:

本文研究规模不变组合半老虎机问题中FTPL策略的最优性与复杂度,证明其在Fréchet和Pareto分布下的遗憾界,并提出条件几何重采样以降低计算复杂度。

AI中文摘要:

本文研究扰动跟随领导者(FTPL)策略在规模不变组合半老虎机问题中的最优性与复杂度。近期,Honda等人(2023)和Lee等人(2024)证明了在具有Fréchet型分布的标准多臂老虎机问题中,FTPL能够实现两全其美(BOBW)最优性。然而,FTPL在组合半老虎机问题中的最优性仍不清楚。本文考虑在规模不变半老虎机设定下带有几何重采样(GR)的FTPL的遗憾界,证明FTPL在Fréchet分布下分别取得$O\left(\sqrt{m^2 d^\frac{1}{\alpha}T}+\sqrt{mdT}\right)$的遗憾,并在对抗性设定下以Pareto分布取得$O\left(\sqrt{mdT}\right)$的最佳可能遗憾界。此外,我们将条件几何重采样(CGR)扩展到规模不变半老虎机设定,将计算复杂度从原始GR的$O(d^2)$降低到$O\left(md\left(\log(d/m)+1\right)\right)$,且不牺牲FTPL的遗憾性能。

英文摘要:

This paper studies the optimality and complexity of Follow-the-Perturbed-Leader (FTPL) policy in size-invariant combinatorial semi-bandit problems. Recently, Honda et al. (2023) and Lee et al. (2024) showed that FTPL achieves Best-of-Both-Worlds (BOBW) optimality in standard multi-armed bandit problems with Fréchet-type distributions. However, the optimality of FTPL in combinatorial semi-bandit problems remains unclear. In this paper, we consider the regret bound of FTPL with geometric resampling (GR) in size-invariant semi-bandit setting, showing that FTPL respectively achieves $O\left(\sqrt{m^2 d^\frac{1}αT}+\sqrt{mdT}\right)$ regret with Fréchet distributions, and the best possible regret bound of $O\left(\sqrt{mdT}\right)$ with Pareto distributions in adversarial setting. Furthermore, we extend the conditional geometric resampling (CGR) to size-invariant semi-bandit setting, which reduces the computational complexity from $O(d^2)$ of original GR to $O\left(md\left(\log(d/m)+1\right)\right)$ without sacrificing the regret performance of FTPL.

补充信息

↑