稀疏线性赌博机中固定预算下的最佳臂识别
Fixed-Budget Best-Arm Identification in Sparse Linear Bandits
- CNRS at CREATE(法国国家科学研究中心驻CREATE)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对稀疏高维线性赌博机的固定预算最佳臂识别问题,提出两阶段算法 Lasso-OD,先利用阈值 Lasso 估计支持集,再在支持集上运行 OD-LinBAI,使错误概率指数仅依赖稀疏度 s 而非维度 d,并证明其几乎极小极大最优。
AI中文摘要:
我们研究固定预算设置下稀疏线性赌博机中的最佳臂识别问题。在稀疏线性赌博机中,未知特征向量 $θ^*$ 可能具有较大的维度 $d$,但其中只有少数特征,例如 $s \ll d$ 个特征具有非零值。我们设计了一个两阶段算法,即基于 Lasso 和最优设计(Lasso-OD)的线性最佳臂识别算法。Lasso-OD 的第一阶段利用特征向量的稀疏性,应用 Zhou (2009) 提出的阈值 Lasso,通过所选臂的奖励和设计矩阵的审慎选择,以高概率正确估计 $θ^*$ 的支持集。Lasso-OD 的第二阶段在该估计的支持集上应用 Yang 和 Tan (2022) 提出的 OD-LinBAI 算法。我们通过仔细选择超参数(如 Lasso 的正则化参数)并平衡两个阶段的错误概率,推导出 Lasso-OD 错误概率的非渐近上界。对于固定的稀疏度 $s$ 和预算 $T$,Lasso-OD 错误概率中的指数依赖于 $s$ 而不依赖于维度 $d$,从而为稀疏且高维的线性赌博机带来显著的性能提升。此外,我们证明 Lasso-OD 在指数意义下几乎达到极小极大最优。最后,我们提供数值示例,证明与 OD-LinBAI、BayesGap、Peace、LinearExploration 和 GSE 等现有非稀疏线性赌博机算法相比,Lasso-OD 具有显著的性能提升。
英文摘要:
We study the best-arm identification problem in sparse linear bandits under the fixed-budget setting. In sparse linear bandits, the unknown feature vector $θ^*$ may be of large dimension $d$, but only a few, say $s \ll d$ of these features have non-zero values. We design a two-phase algorithm, Lasso and Optimal-Design- (Lasso-OD) based linear best-arm identification. The first phase of Lasso-OD leverages the sparsity of the feature vector by applying the thresholded Lasso introduced by Zhou (2009), which estimates the support of $θ^*$ correctly with high probability using rewards from the selected arms and a judicious choice of the design matrix. The second phase of Lasso-OD applies the OD-LinBAI algorithm by Yang and Tan (2022) on that estimated support. We derive a non-asymptotic upper bound on the error probability of Lasso-OD by carefully choosing hyperparameters (such as Lasso's regularization parameter) and balancing the error probabilities of both phases. For fixed sparsity $s$ and budget $T$, the exponent in the error probability of Lasso-OD depends on $s$ but not on the dimension $d$, yielding a significant performance improvement for sparse and high-dimensional linear bandits. Furthermore, we show that Lasso-OD is almost minimax optimal in the exponent. Finally, we provide numerical examples to demonstrate the significant performance improvement over the existing algorithms for non-sparse linear bandits such as OD-LinBAI, BayesGap, Peace, LinearExploration, and GSE.