arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2403.01315cs.LGstat.ML

Sleeping Bandits的近最优逐动作遗憾界

Near-optimal Per-Action Regret Bounds for Sleeping Bandits

  • University of Victoria(维多利亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Quan Nguyen, Nishant A. Mehta

更新

AI总结:

针对每轮可用臂及损失由对手决定的sleeping bandits问题,通过广义EXP3、EXP3-IX和带Tsallis熵的FTRL直接最小化逐动作遗憾,推导出近最优遗憾界,并扩展至带sleeping experts建议的设定及置信度报告场景。

AI中文摘要:

我们推导了sleeping bandits的近最优逐动作遗憾界,其中每轮可用的臂集及其损失均由对手选择。在共有$K$个臂、每轮最多$A$个可用臂、共$T$轮的设定下,已知的最佳上界为$O(K\sqrt{TA\ln{K}})$,该上界是通过间接最小化内部sleeping regret获得的。与极小化$Ω(\sqrt{TA})$下界相比,该上界包含一个额外的乘法因子$K\ln{K}$。我们通过使用EXP3、EXP3-IX和带有Tsallis熵的FTRL的广义版本直接最小化逐动作遗憾来解决这一差距,从而获得了阶为$O(\sqrt{TA\ln{K}})$和$O(\sqrt{T\sqrt{AK}})$的近最优界。我们将结果扩展到带有sleeping experts建议的bandits设定,并在此过程中推广了EXP4。这为标准非sleeping bandits的许多现有自适应和跟踪遗憾界提供了新证明。将结果扩展到报告其置信度的experts的bandit版本,得出了主要取决于experts置信度总和的置信度遗憾新界。我们证明了一个下界,表明对于任何极小化最优算法,都存在一个动作,其遗憾在$T$中是次线性的,但在其活跃轮数中是线性的。

英文摘要:

We derive near-optimal per-action regret bounds for sleeping bandits, in which both the sets of available arms and their losses in every round are chosen by an adversary. In a setting with $K$ total arms and at most $A$ available arms in each round over $T$ rounds, the best known upper bound is $O(K\sqrt{TA\ln{K}})$, obtained indirectly via minimizing internal sleeping regrets. Compared to the minimax $Ω(\sqrt{TA})$ lower bound, this upper bound contains an extra multiplicative factor of $K\ln{K}$. We address this gap by directly minimizing the per-action regret using generalized versions of EXP3, EXP3-IX and FTRL with Tsallis entropy, thereby obtaining near-optimal bounds of order $O(\sqrt{TA\ln{K}})$ and $O(\sqrt{T\sqrt{AK}})$. We extend our results to the setting of bandits with advice from sleeping experts, generalizing EXP4 along the way. This leads to new proofs for a number of existing adaptive and tracking regret bounds for standard non-sleeping bandits. Extending our results to the bandit version of experts that report their confidences leads to new bounds for the confidence regret that depends primarily on the sum of experts' confidences. We prove a lower bound, showing that for any minimax optimal algorithms, there exists an action whose regret is sublinear in $T$ but linear in the number of its active rounds.

补充信息

↑