发表机构
Peking University; Stanford University; Shanghai Jiao Tong University(北京大学; 斯坦福大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多臂老虎机场景,证明了后悔与不稳定性的乘积下界,提出SLE-UCB算法匹配该下界,解决了相关开放问题。
AI 中文摘要
多臂老虎机算法通过后悔值进行评估,但相同的后悔值可能在独立运行中对应不同的分配情况。针对K个臂、T轮的场景,我们研究最坏情况后悔值$\boldsymbol{\textit{R}}_{K,T}$与不稳定性$\boldsymbol{\textit{S}}_{K,T}$的权衡,其中不稳定性定义为最终拉杆次数的最大标准差。在有限时间后悔条件下,且不采用以往渐近分析中的正则性假设,我们证明有限时间下界$\boldsymbol{\textit{R}}_{K,T}\boldsymbol{\textit{S}}_{K,T}\boldsymbol{\boldsymbol{\text{≥}}}\boldsymbol{C}\boldsymbol{T}^{3/2}$,其中常数C与K、T无关。我们还提出了一种新的可调算法Stabilized Lower-Envelope UCB(简称SLE-UCB),该算法结合了运行下界包络指标与递减拉杆次数稳定器。SLE-UCB满足$\boldsymbol{\textit{R}}_{K,T}\boldsymbol{\textit{S}}_{K,T}\boldsymbol{=}\boldsymbol{O}(\boldsymbol{T}^{3/2}\boldsymbol{\text{log}}\boldsymbol{K})$,其隐含常数与K、T无关,在T维度上与下界完全匹配,在K维度上仅相差一个对数因子。为证明不稳定性边界,我们开发了一种新的离线顶部前缀表示法,消除了在线决策的路径依赖,结合单奖励扰动与Efron–Stein不等式,该表示法可控制拉杆次数的方差。因此,后悔值与不稳定性呈关于K的倒数关系,而它们的乘积对K无多项式依赖。这些结果解决了文献中提出的关于与臂相关的精确后悔-不稳定性前沿的开放问题。
英文摘要
Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-case regret $\mathcal{R}_{K,T}$ and instability $\mathcal S_{K,T}$, defined as the largest standard deviation of a terminal pull count, for $K$ arms and $T$ rounds. We prove the finite-time lower bound $\mathcal R_{K,T}\mathcal S_{K,T}\ge C T^{3/2}$, where $C$ is independent of $K$ and $T$, under a finite-time regret condition and without the regularity assumptions imposed in the prior asymptotic analysis. We also introduce Stabilized Lower-Envelope UCB (\textup{\textsc{SLE-UCB}}), a new tunable algorithm combining a running lower-envelope index with a decreasing pull-count stabilizer. \textup{\textsc{SLE-UCB}} satisfies $\mathcal R_{K,T}\mathcal S_{K,T}=O(T^{3/2}\log K)$, with an implicit constant independent of $K$ and $T$, matching the lower bound exactly in $T$ and within a logarithmic factor in $K$. To prove the instability bound, we develop a new offline top-prefix representation that removes path dependence from online decisions. Together with single-reward perturbations and the Efron--Stein inequality, this representation controls pull-count variance. Thus, regret and instability depend reciprocally on $K$, while their product has no polynomial dependence on $K$. These results resolve the open question raised in the literature concerning the sharp arm-dependent regret--instability frontier.