arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2105.14267stat.MLcs.LGmath.STstat.TH

Information Directed Sampling for Sparse Linear Bandits

  • Deepmind(深度思维)
  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

Botao Hao, Tor Lattimore, Wei Deng

更新

英文摘要:

Stochastic sparse linear bandits offer a practical model for high-dimensional online decision-making problems and have a rich information-regret structure. In this work we explore the use of information-directed sampling (IDS), which naturally balances the information-regret trade-off. We develop a class of information-theoretic Bayesian regret bounds that nearly match existing lower bounds on a variety of problem instances, demonstrating the adaptivity of IDS. To efficiently implement sparse IDS, we propose an empirical Bayesian approach for sparse posterior sampling using a spike-and-slab Gaussian-Laplace prior. Numerical results demonstrate significant regret reductions by sparse IDS relative to several baselines.

↑