具有任意自适应动作集的线性上下文赌博机的近极小极大最优遗憾
Nearly Minimax-Optimal Regret for Linear Contextual Bandits with Arbitrary Adaptive Action Sets
浏览论文内容
中文总结 AI 辅助
研究具有任意自适应动作集的线性上下文赌博机,建立了匹配的上下界,改进了对动作集大小的依赖,并证明多项式依赖最优。
中文摘要 AI 辅助
我们研究具有任意动作菜单的随机线性上下文赌博机,这些动作菜单可能依赖于固定参数和交互历史。我们建立了匹配的上下界,直至对数因子。设$d$为维度,$K$为菜单大小,$T$为时间范围。对于$2\le K\le d$,我们证明了一个上界$\widetilde O(K^{1/4}\sqrt{dT})$。当$T\ge d^2$时,我们进一步证明了一个下界$\Omega(K^{1/4}\sqrt{dT})$。因此,对于$T\ge d^2$且$2\le K\le d$,上下界在对数因子内匹配,且对$K$的多项式依赖是最优的。与之前的$\widetilde O(\sqrt{dKT})$界相比,我们的上界将$K$的依赖改进了$K^{1/4}$倍。对于$K\ge d$,我们证明了一个上界$\widetilde O_{d,T}\left(\sqrt{dT}\min\{\sqrt d,(d\log K)^{1/4}\}\right)$和一个下界$\Omega\left(\sqrt{dT}\min\left\{\sqrt d,\left(\frac{d\log K}{\log(2d)}\right)^{1/4}\right\}\right)$。这里,$\widetilde O_{d,T}$仅省略了$d$和$T$中的对数因子。特别地,对于多项式大的$K\ge d$,上下界均按$d^{3/4}\sqrt T$缩放(直至对数因子),比标准$\widetilde O(d\sqrt T)$速率改进了$d^{1/4}$倍。随着$K$进一步增长,一旦$\log K$达到$d$量级,遗憾平滑恢复到$d\sqrt T$尺度。
英文摘要
We study stochastic linear contextual bandits with arbitrary action menus that may depend on the fixed parameter and the interaction history. We establish matching upper and lower bounds, up to logarithmic factors. Let $d$ be the dimension, $K$ be the menu size, and $T$ the time horizon. For $2\le K\le d$, we prove an upper bound $\widetilde O(K^{1/4}\sqrt{dT})$. When $T\ge d^2$, we further prove a lower bound $Ω(K^{1/4}\sqrt{dT})$. Thus, for $T\ge d^2$ and $2\le K\le d$, the upper and lower bounds match up to logarithmic factors, and the polynomial dependence on $K$ is optimal. Compared with the previous $\widetilde O(\sqrt{dKT})$ bound, our upper bound improves the dependence on $K$ by a factor of $K^{1/4}$. For $K\ge d$, we prove an upper bound $\widetilde O_{d,T}\left(\sqrt{dT}\min\{\sqrt d,(d\log K)^{1/4}\}\right)$ and a lower bound $Ω\left(\sqrt{dT}\min\left\{\sqrt d,\left(\frac{d\log K}{\log(2d)}\right)^{1/4}\right\}\right)$. Here, $\widetilde O_{d,T}$ omits logarithmic factors only in $d$ and $T$. In particular, for polynomially large $K\ge d$, the upper and lower bounds both scale as $d^{3/4}\sqrt T$ up to logarithmic factors, improving the standard $\widetilde O(d\sqrt T)$ rate by a factor of $d^{1/4}$. As $K$ grows further, the regret smoothly recovers the $d\sqrt T$ scale once $\log K$ reaches order $d$.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。