arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.07307cs.LGstat.ML

Follow-the-Perturbed-Leader 方法在 m-集半臂赌博机问题中实现双世界最优

Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems

  • School of Mathematical Sciences, Peking University(北京大学数学科学学院)

机构由 AI 辅助整理,请以论文原文为准。

Jingxin Zhan, Yuchen Xin, Chenjie Sun, Zhihua Zhang

更新

AI总结:

针对m-集半臂赌博机问题,提出带Fréchet扰动的Follow-the-Perturbed-Leader(FTPL)策略,在对抗性设置中实现接近最优遗憾界O(√nm(√d log(d)+m^{5/6})),并在随机设置中获得对数遗憾,且无需显式优化采样概率。

AI中文摘要:

我们考虑组合半臂赌博机问题的一个常见情况,即m-集半臂赌博机,学习者从总共d个臂中精确选择m个臂。在对抗性设置中,已知的最佳遗憾界为时间范围n下的O(√nmd),由著名的Follow-the-Regularized-Leader(FTRL)策略实现。然而,这需要在每个时间步通过优化问题显式计算臂选择概率,并根据这些概率进行采样。Follow-the-Perturbed-Leader(FTPL)策略可避免此问题,它只需选择在随机扰动下排名前m的臂。本文表明,在对抗性设置中,带有Fréchet扰动的FTPL也能获得接近最优的遗憾界O(√nm(√d log(d)+m^{5/6})),并实现双世界最优遗憾界,即在随机设置中获得对数遗憾。此外,我们的下界表明,使用我们的方法无法避免额外因子;任何改进都需要根本不同且更具挑战性的方法。

英文摘要:

We consider a common case of the combinatorial semi-bandit problem, the $m$-set semi-bandit, where the learner exactly selects $m$ arms from the total $d$ arms. In the adversarial setting, the best regret bound, known to be $\mathcal{O}(\sqrt{nmd})$ for time horizon $n$, is achieved by the well-known Follow-the-Regularized-Leader (FTRL) policy. However, this requires to explicitly compute the arm-selection probabilities via optimizing problems at each time step and sample according to them. This problem can be avoided by the Follow-the-Perturbed-Leader (FTPL) policy, which simply pulls the $m$ arms that rank among the $m$ smallest (estimated) loss with random perturbation. In this paper, we show that FTPL with a Fréchet perturbation also enjoys the near optimal regret bound $\mathcal{O}(\sqrt{nm}(\sqrt{d\log(d)}+m^{5/6}))$ in the adversarial setting and approaches best-of-both-world regret bounds, i.e., achieves a logarithmic regret for the stochastic setting. Moreover, our lower bounds show that the extra factors are unavoidable with our approach; any improvement would require a fundamentally different and more challenging method.

↑