后验回火解释线性与广义线性汤普森采样中的方差膨胀
Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling
- Daniels School of Business, Purdue University(普渡大学丹尼尔斯商学院)
- Department of Statistics, University of Wisconsin-Madison(威斯康星大学麦迪逊分校统计系)
- Department of Statistics, Texas A&M University(德克萨斯农工大学统计系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出α-汤普森采样算法,确定先验与奖励分布的正则性条件,推导其遗憾界并解释线性和广义线性汤普森采样中d^(3/2)因子的来源。
AI中文摘要:
我们研究了一种名为α-汤普森采样(α-TS)的汤普森采样(TS)算法变体,用于求解随机广义线性多臂老虎机问题。现有TS分析需要对方差进行膨胀以推导近最优的遗憾保证,我们通过引入使用分数后验或α-后验而非标准后验的α-TS,将方差膨胀的思路形式化。我们的主要贡献是确定了先验和奖励分布上的一般正则性条件,这些条件使我们能够对α-TS进行遗憾分析,且无需像以往工作那样假设后验分布存在任何易处理的近似。对于α∝d⁻¹的特定选择,我们的一般遗憾界在指数族和次高斯族奖励分布下均得到了已知最优的O(d^(3/2)√T log T)遗憾界。我们进一步提供了一个依赖于α的下界,表明遗憾常数依赖于乘积αd,且当α∝d⁻¹时,遗憾按Ω(d^(3/2)√T)缩放,解释了上界中d^(3/2)因子的来源。我们的证明技术改编并结合了线性多臂老虎机问题分析的最新进展与贝叶斯统计文献中的一阶和二阶后验集中理论。
英文摘要:
We study a variant of the Thompson Sampling (TS) algorithm, called $α$-TS, for solving stochastic generalized linear bandit problems. Existing analyses of TS require inflating the posterior variance to derive near-optimal regret guarantees. We formalize the idea of variance inflation by introducing $α$-TS that uses a fractional or $α$-posterior instead of the standard posterior. Our main contribution is to identify general regularity conditions on the prior and reward distributions that enable a regret analysis of $α$-TS without assuming any tractable approximation of the posterior distribution, unlike previous works. For a specific choice of $α\propto d^{-1}$, our general regret bound yields the best known regret bound of $O(d^{3/2}\sqrt{T}\log T)$ for both the exponential and sub-Gaussian families of reward distributions. We further provide an $α$-dependent lower bound showing that the regret constant depends on the product $αd$, and that when $α\propto d^{-1}$ the regret scales as $Ω(d^{3/2}\sqrt{T})$, explaining the origin of the $d^{3/2}$ factor in the upper bound. Our proof technique adapts and combines recent advancements in the analysis of linear bandit problems with first- and second-order posterior concentration theory from the Bayesian statistics literature.