发表机构
University of California, Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究具有多个最优臂的多臂老虎机,提出了改进的极小极大遗憾界并匹配下界,表明对最优臂数量的近似了解对实现接近最优性能是必要的。
AI 中文摘要
我们研究具有多个最优臂的多臂老虎机(MAB),其动机是许多实际决策问题允许多个正确答案。对于具有$A$个最优臂的$K$臂老虎机,我们首先对先前的子采样算法(De Heide等人,2021;Zhu和Nowak,2020)提供了更尖锐的分析,建立了$\tilde{O}\big(\frac{K-A}{\sqrt{KA}}\sqrt{T}\big)$的极小极大遗憾,其中$T$是总交互次数,$\tilde O(\\.cdot)$忽略所有常数和对数因子,改进了之前的$\tilde{O}(\\.sqrt{KT/A})$遗憾。然后我们提供了匹配的下界(至多相差对数因子),表明我们建立的速率几乎是极小极大最优的。我们进一步表明,对$A$的了解(至多相差$\tilde{O}(1)$因子)对于实现接近最优的遗憾是必要的,因为针对一个最优臂数量的接近最优算法,在最优臂数量较小时,必须比最优遗憾产生显著更大的遗憾。总体而言,我们的结果为$1 \leq A \leq K-1$整个范围内的$K$臂老虎机提供了全面的极小极大特征描述。
英文摘要
We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 2021; Zhu and Nowak, 2020), establishing a $\tilde{O}\Big(\frac{K-A}{\sqrt{KA}}\sqrt{T} \Big)$ minimax regret, where $T$ is the total number of interactions and $\tilde O(\cdot)$ drops all constant and logarithmic factors, improving the previous $\tilde{O}(\sqrt{KT/A})$ regret. We then provide a matching lower bound up to logarithmic factors, indicating that our established rate is nearly minimax-optimal. We further show that the knowledge of $A$ up to $\tilde{O}(1)$ factors is necessary to achieve near-optimal regret, as near-optimal algorithms for one number of optimal arms must incur substantially larger regret than optimal regret for a smaller number. Overall, our results provide a comprehensive minimax characterization of $K$-armed bandits with $A$ over the entire range of $1 \leq A \leq K-1$.