PopArt:用于最优稀疏线性Bandit的高效稀疏回归与实验设计
PopArt: Efficient Sparse Regression and Experimental Design for Optimal Sparse Linear Bandits
- University of Arizona(亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出高效稀疏线性估计方法PopArt及凸实验设计准则,推导出改进regret上界的稀疏线性bandit算法,并在数据匮乏场景下证明了匹配下界。
AI中文摘要:
在稀疏线性bandit中,学习智能体顺序地选择动作并接收奖励反馈,且奖励函数线性依赖于动作协变量的少数几个坐标。这在许多现实世界的序贯决策问题中具有应用。在本文中,我们提出了一种简单且计算高效的稀疏线性估计方法,称为PopArt,与Lasso(Tibshirani, 1996)相比,它在许多问题中享有更紧的$\ell_1$恢复保证。我们的界自然地激发了一种实验设计准则,该准则是凸的,因此在计算上易于求解。基于我们新颖的估计器和设计准则,我们推导出了稀疏线性bandit算法,该算法享有优于现有技术(Hao et al., 2020)的regret上界,特别是关于给定动作集的几何形状。最后,我们在数据匮乏的情况下证明了稀疏线性bandit的匹配下界,这弥补了先前工作中上界与下界之间的差距。
英文摘要:
In sparse linear bandits, a learning agent sequentially selects an action and receive reward feedback, and the reward function depends linearly on a few coordinates of the covariates of the actions. This has applications in many real-world sequential decision making problems. In this paper, we propose a simple and computationally efficient sparse linear estimation method called PopArt that enjoys a tighter $\ell_1$ recovery guarantee compared to Lasso (Tibshirani, 1996) in many problems. Our bound naturally motivates an experimental design criterion that is convex and thus computationally efficient to solve. Based on our novel estimator and design criterion, we derive sparse linear bandit algorithms that enjoy improved regret upper bounds upon the state of the art (Hao et al., 2020), especially w.r.t. the geometry of the given action set. Finally, we prove a matching lower bound for sparse linear bandits in the data-poor regime, which closes the gap between upper and lower bounds in prior work.