1/2-Tsallis-INF 是否也能很好地用于最优臂识别?
Does 1/2-Tsallis-INF Also Work Well for Best-Arm Identification?
浏览论文内容
中文总结 AI 辅助
该研究探讨 1/2-Tsallis-INF 能否用于最优臂识别,通过构造 Lyapunov 函数得出其失败概率的多项式上下界,证明指数 2 本质紧。
中文摘要 AI 辅助
遗憾最小化(RM)和最优臂识别(BAI)是多臂老虎机的两个基本目标。在遗憾最小化算法中,1/2-Tsallis-INF 是一种典型的两全其美的 FTRL(跟随正则化领导者)算法:它在随机老虎机中实现对数伪遗憾,同时在对抗老虎机中保持极小极大最优遗憾,且无需预先了解环境。这引发了一个自然问题:该算法在无需额外探索的情况下,是否也能可靠地识别最优臂?我们通过分析失败概率 Err_t 来研究随机老虎机中的这一问题,Err_t 定义为 1/2-Tsallis-INF 的累积重要性加权损失估计所确定的经验最优臂与真实最优臂不同的概率。主要难点在于,在对数遗憾尺度下,次优臂的采样概率近似为 1/t 量级。因此,重要性加权导致累积估计量的波动幅度与其均值分离度处于同一线性尺度。为克服这一障碍,我们在扩散玩具模型的指导下,为最优臂与最佳竞争臂的估计累积损失之间的间隙过程构造了一个 Lyapunov 函数。这得出了 Err_t 的多项式上界:对于学习率 η_t=α/√t,Err_t 以 t^{-2+α²μ_{i_*}/4+ρ} 的速率衰减,其中 μ_{i_*} 为真实最优臂的平均损失,ρ 为任意正实数。我们还建立了任意 ε>0 时的下界 Ω(t^{-2-ε}),表明指数 2 本质上是紧的。
英文摘要
Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits. Among regret-minimizing algorithms, $1/2$-Tsallis-INF is a canonical best-of-both-worlds FTRL algorithm: it achieves logarithmic pseudo-regret in stochastic bandits while retaining minimax-optimal regret in adversarial bandits, without knowing the environment in advance. This raises a natural question: can the same algorithm, without additional exploration, also identify the best arm reliably? We study this question in stochastic bandits by analyzing the failure probability $\operatorname{Err}_t$, defined as the probability that the empirical best arm determined by the cumulative importance-weighted loss estimates of 1/2-Tsallis-INF differs from the true optimal arm. The main difficulty is that, at the logarithmic-regret scale, suboptimal arms are sampled with probability heuristically of order $1/t$. Consequently, importance weighting causes the cumulative estimator to fluctuate on the same linear scale as its mean separation. To overcome this obstacle, guided by a diffusion toy model, we construct a Lyapunov function for the gap process between the estimated cumulative loss of the optimal arm and that of the best competing arm. This leads to polynomial upper bounds on $\operatorname{Err}_t$: for learning rate $η_t=α/\sqrt t$, $\operatorname{Err}_t$ decays at rate $t^{-2+α^2μ_{i_*}/4+ρ}$ for any $ρ>0$, where $μ_{i_*}$ denotes the mean loss of the true optimal arm. We also establish a lower bound $Ω(t^{-2-\varepsilon})$ for any $\varepsilon>0$, showing that the exponent $2$ is essentially tight.
发表机构
- School of Mathematical Sciences, Peking University(北京大学数学科学学院)
- Center for Applied Statistics and School of Statistics, Renmin University of China(中国人民大学应用统计中心与统计学院)
机构由 AI 辅助整理,请以论文原文为准。