arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2211.05964stat.MLcs.LGmath.STstat.MEstat.TH

Thompson Sampling for High-Dimensional Sparse Linear Contextual Bandits

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

Sunrit Chakraborty, Saptarshi Roy, Ambuj Tewari

更新

英文摘要:

We consider the stochastic linear contextual bandit problem with high-dimensional features. We analyze the Thompson sampling algorithm using special classes of sparsity-inducing priors (e.g., spike-and-slab) to model the unknown parameter and provide a nearly optimal upper bound on the expected cumulative regret. To the best of our knowledge, this is the first work that provides theoretical guarantees of Thompson sampling in high-dimensional and sparse contextual bandits. For faster computation, we use variational inference instead of Markov Chain Monte Carlo (MCMC) to approximate the posterior distribution. Extensive simulations demonstrate the improved performance of our proposed algorithm over existing ones.

补充信息

↑