发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对核函数多臂老虎机现有下界仅适用于特定核的问题,本文建立了非恒定连续核的通用极小极大后悔下界,明确了其近最优性及对数因子的不可避免性,还推导了特定核的最优缩放规律。
AI 中文摘要
核函数老虎机问题是指在给定再生核希尔伯特空间(RKHS)中,以带噪反馈依次优化一个范数有界的未知函数。核函数老虎机后悔分析中的核心量是最大信息增益$\gamma_T$。具体而言,现有最优上界的缩放形式为$\sqrt{T\gamma_T}$(含对数因子),且已针对平方指数核、马顿核等特定核推导了几乎匹配的下界。然而,通用核的下界仍缺失,导致尚不清楚上界在何种通用性下接近最优。本文针对紧域上的非恒定连续核,建立了通用$\Omega(\sqrt{T\gamma_T/\log T})$极小极大后悔下界,在极广范围内确定了其近最优性(对数因子范围内)。我们表明,该下界中出现的对数因子在一般情况下不可避免,但在特定条件下可消除。此外,我们的结果意味着,对于满足$\nu \in (0,2)$的马顿-$\nu$核、满足$\gamma \in (0,2)$的$\gamma$指数核及某些分段多项式核,极小极大最优缩放恰好为$\Theta(\sqrt{T\gamma_T})$(即常数因子范围内)。
英文摘要
The kernel bandit problem consists of sequentially optimizing an unknown function with noisy feedback, where the function has bounded norm in a given Reproducing Kernel Hilbert Space (RKHS). A central quantity in the regret analysis of kernel bandits is the maximum information gain $γ_T$. In particular, the best existing upper bounds scale as $\sqrt{Tγ_T}$ up to log factors, and nearly-matching lower bounds have been derived for specific kernels such as squared exponential and Matérn. However, lower bounds for general kernels are lacking, thus making it unclear in what generality the upper bounds are near-optimal. In this paper, we establish a general $Ω(\sqrt{Tγ_T/\log T})$ minimax regret lower bound for non-constant continuous kernels on compact domains, establishing near-optimality (within log factors) in a very general sense. We show that the log factor appearing in this bound is unavoidable in general, but that it can be removed under certain conditions. Among other things, our findings imply that the minimax-optimal scaling is exactly $Θ(\sqrt{Tγ_T})$ (i.e., within constant factors) for the Matérn-$ν$ kernel with $ν\in (0,2)$, $γ$-exponential kernel with $γ\in (0,2)$, and certain piecewise-polynomial kernels.