改进型多臂老虎机问题的无先验竞争比:尺度、曲率与时间范围免费,但在噪声下无法同时获得
Prior-Free Competitive Ratios for Improving Bandits: Scale, Curvature and Horizon Are Free, but Not Jointly Under Noise
浏览论文内容
中文总结 AI 辅助
本文研究改进型多臂老虎机问题,提出无需先验的探测-提交算法,实现最优竞争比 $\Theta(\sqrt k+k/T)$,并揭示噪声下先验代价的跃升。
中文摘要 AI 辅助
在改进型多臂老虎机问题中,每个臂 $i$ 具有未知的非递减、离散凹奖励曲线 $f_i$,第 $t$ 次拉动臂 $i$ 获得奖励 $f_i(t)$。对于足够长的时间范围,Blum 和 Ravichandran (ALT 2025) 证明了当最优臂的尺度 $m=f^*(T)$ 已知时($T\ge2k$),随机算法可以达到最优单臂的 $O(\sqrt k)$ 近似;当尺度未知时($T>4k$),近似比为 $O(\sqrt k\log k)$,而对应的下界为 $\Omega(\sqrt k)$。对数因子并非必要:一个仅一页的“探测-提交”算法在 $T\ge2\lfloor\sqrt k\rfloor$ 时,无需任何尺度知识即可达到竞争比 $4\sqrt3\\,\sqrt k$,并且我们确定了每个时间范围的最优竞争比 $\Theta(\sqrt k+k/T)$,该结果同样适用于未知时间范围。在没有噪声的情况下,完全不需要任何先验:一个随机边际探测算法既不读取尺度 $m$,也不读取 Blum、Garicano、Ravichandran 和 Sharma (UAI 2026) 提出的凹包络指数 $\beta$,更不读取时间范围 $T$,却能同时对每个 $\beta$ 和每个时间范围达到最优的 $\Theta(k^{\beta/(1+\beta)}+k/T)$。在 Blum 和 Ravichandran 的乘性噪声模型下,探测-提交算法在不知道噪声水平的情况下保持相同的全时间范围阶 $\Theta(\sqrt k+k/T)$(在同一范围内为 $\Theta(\sqrt k)$),但先验的代价急剧上升:对于任意固定噪声水平 $\varepsilon\in(0,1/2]$,均匀适应代价 $\phi_\varepsilon(k)$——即对于既不知道 $m$ 也不知道 $\beta$ 的算法,在时间范围 $T\ge16k$ 上相对于 $k^{\beta/(1+\beta)}$ 的损失的最坏情况——为 $\Theta_\varepsilon(\sqrt{\log k/\log\log k})$,该下界在固定正 $\varepsilon$ 下关于 $k$ 渐近成立,并由一个嵌套随机排列探测算法匹配;而仅知道 $m$ 或 $\beta$ 之一即可恢复常数代价。
英文摘要
In the improving multi-armed bandits problem, each of $k$ arms has an unknown nondecreasing, discretely concave reward curve $f_i$, and pulling arm $i$ for the $t$-th time yields $f_i(t)$. For sufficiently long horizons, Blum and Ravichandran (ALT 2025) proved that randomized algorithms achieve an $O(\sqrt k)$ approximation to the best single arm when the scale $m=f^*(T)$ of the optimal arm is known ($T\ge2k$), and $O(\sqrt k\log k)$ when it is not ($T>4k$), against an $Ω(\sqrt k)$ lower bound. The logarithmic factor is unnecessary: a one-page \emph{probe-and-commit} algorithm achieves competitive ratio $4\sqrt3\,\sqrt k$ for $T\ge2\lfloor\sqrt k\rfloor$, without any knowledge of the scale, and we determine the optimal ratio for every horizon, $Θ(\sqrt k+k/T)$, also for unknown horizons. Without noise, \emph{no prior is needed at all}: a random-marginal probing algorithm reading neither the scale $m$, nor the concavity-envelope exponent $β$ of Blum, Garicano, Ravichandran and Sharma (UAI 2026), nor the horizon $T$, achieves the optimal $Θ(k^{β/(1+β)}+k/T)$ simultaneously for every $β$ and every horizon. Under the multiplicative noise model of Blum and Ravichandran, probe-and-commit keeps the same all-horizon order $Θ(\sqrt k+k/T)$ without knowing the noise level (and $Θ(\sqrt k)$ on the same range), but the price of priors jumps: for any fixed noise level $\varepsilon\in(0,1/2]$, the uniform price of adaptation $ϕ_\varepsilon(k)$ --- the worst case over horizons $T\ge16k$ of the loss relative to $k^{β/(1+β)}$ for algorithms knowing neither $m$ nor $β$ --- is $Θ_\varepsilon(\sqrt{\log k/\log\log k})$, the lower bound asymptotic in $k$ at fixed positive $\varepsilon$ and matched by a nested random-permutation probing algorithm, whereas knowing either $m$ or $β$ alone restores a constant price.
发表机构
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。