arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

打破总方差障碍:具有固定动作集的线性异方差博弈的精确样本复杂度

Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set

Heyang Zhao, Tianyuan Jin, Weixin Wang, Vincent Y. F. Tan, Pan Xu, Quanquan Gu

arXiv 2607.23679首次发表:更新:

发表机构

University of California, Los Angeles; National University of Singapore; Duke University(加利福尼亚大学洛杉矶分校; 新加坡国立大学; 杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究具有固定动作集的线性异方差博弈问题,提出方差自适应算法\texttt{VAEE}及方差感知变体,实现了具有近乎调和均值依赖率的简单遗憾,打破了$\sqrt{\Lambda}$障碍,并建立了近乎匹配的下界。

AI 中文摘要

近年来,在博弈和强化学习中处理异方差噪声的兴趣日益增加。在这些工作中,噪声的累积方差$\Lambda = \sum_{t=1}^T \sigma_t^2$(其中$\sigma_t^2$是第$t$轮噪声的方差)用于表征问题的统计复杂度,对于具有异方差噪声的$d$维线性博弈,产生了$\tilde{\cal{O}}(d \sqrt{\Lambda / T^2})$阶的简单遗憾界。然而,仔细观察会发现,即使在一半轮次中噪声接近零,$\Lambda$仍保持相同阶数,这表明对$\Lambda$的依赖不是最优的。本文重新审视了具有异方差噪声且动作集在整个学习过程中固定的随机线性博弈问题。对于大动作集,我们提出了一种新颖的方差自适应算法\texttt{VAEE}(带消除的方差感知探索),它在未被消除的候选动作集中积极探索使信息增益最大化的动作。通过主动探索策略,我们表明\texttt{VAEE}实现了具有近乎调和均值依赖率的简单遗憾。对于有限多个动作,我们提出了基于$G$最优设计探索的方差感知变体,它实现了对$d$具有更精确依赖的简单遗憾。我们还为固定动作集设置建立了一个近乎匹配的下界,表明调和均值依赖率是不可避免的。据我们所知,这是第一项打破具有异方差噪声的随机线性博弈的$\sqrt{\Lambda}$障碍的工作。

英文摘要

Recent years have witnessed increasing interests in tackling heteroscedastic noise in bandits and reinforcement learning. In these works, the cumulative variance of the noise $Λ= \sum_{t=1}^T σ_t^2$, where $σ_t^2$ is the variance of the noise at round $t$, is used to characterize the statistical complexity of the problem, yielding \emph{simple regret} bounds of order $\tilde{\cal{O}}(d \sqrt{Λ/ T^2})$ for $d$-dimensional linear bandits with heteroscedastic noise. However, with a closer look, $Λ$ remains the same order even if the noise is close to zero at half of the rounds, which indicates that the $Λ$-dependence is not optimal. In this paper, we revisit the stochastic linear bandit problem with heteroscedastic noise, where the action set is prefixed throughout the learning process. We propose a novel variance-adaptive algorithm \texttt{VAEE} (Variance-Aware Exploration with Elimination) for large action set, which actively explores actions that maximizes the information gain among a candidate set of actions that are not eliminated. With the active-exploration strategy, we show that \texttt{VAEE} achieves a \emph{simple regret} with a nearly \emph{harmonic-mean} dependent rate. For finitely many actions, we propose a variance-aware variant of G-optimal design based exploration, which achieves a simple regret with sharper dependence on $d$. We also establish a nearly matching lower bound for the fixed action set setting indicating that \emph{harmonic-mean} dependent rate is unavoidable. To the best of our knowledge, this is the first work that breaks the $\sqrtΛ$ barrier for stochastic linear bandits with heteroscedastic noise.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑