线性集成采样与更小的集成规模
Linear Ensemble Sampling with Smaller Ensembles
浏览论文内容
中文总结 AI 辅助
针对线性bandit中集成采样规模与遗憾保证的间隙,提出按格拉姆矩阵变化刷新集成的算法,将规模降至Θ(d log d+d log log T),同时保持最优遗憾界,并适用于有限臂集。
中文摘要 AI 辅助
集成采样通过维护一组模型为随机化探索提供了一种实用方法,但集成规模可以多小同时仍保持强遗憾保证仍未解决。特别是,现有保证使用大小为 $\Theta(d\log T)$ 的集成,相对于内在的 $\Omega(d)$ 集成规模障碍,在时间范围 $T$ 上留下了一个对数间隙。我们旨在通过提出一种集成采样算法来缩小这一间隙,该算法仅在正则化格拉姆矩阵发生显著变化时刷新集成。这一机制将扰动分析局部化到具有受控格拉姆矩阵漂移的时期,并将足够的集成规模减少到 $\Theta(d\log d+d\log\log T)$,同时保持集成采样在任意有界臂集上的最先进遗憾 $\tilde O(d^{3/2}\sqrt T)$。我们进一步表明,当臂集是基数为 $K$ 的有限集时,所提出的算法实现了更尖锐的遗憾界 $\tilde O(d\sqrt{T\log K})$。据我们所知,这是第一个同时恢复随机线性bandit算法已知的两种典型遗憾缩放关系的集成采样保证:任意有界臂集的 $\tilde O(d^{3/2}\sqrt{T})$ 速率和有限臂集的 $\tilde O (d\sqrt{T\log K})$ 速率。该算法还支持无需重置过去数据的随时实现,实验表明,在使用显著更小的集成时,它仍与基线保持竞争力。
英文摘要
Ensemble sampling offers a practical approach to randomized exploration by maintaining a collection of models, but how small an ensemble can be while retaining strong regret guarantees remains unresolved. In particular, the existing guarantees use an ensemble size of $Θ(d\log T)$, leaving a logarithmic gap in the horizon $T$ relative to the intrinsic $Ω(d)$ ensemble-size barrier. We aim to narrow this gap by proposing an ensemble sampling algorithm that refreshes the ensemble only when the regularized Gram matrix changes substantially. This mechanism localizes the perturbation analysis to epochs with controlled Gram-matrix drift and reduces the sufficient ensemble size to $Θ(d\log d+d\log\log T)$, while preserving the state-of-the-art $\tilde O(d^{3/2}\sqrt T)$ regret for ensemble sampling with arbitrary bounded arm sets. We further show that, when the arm set is finite of cardinality $K$, the proposed algorithm achieves the sharper regret bound $\tilde O(d\sqrt{T\log K})$. To the best of our knowledge, this is the first ensemble-sampling guarantee that simultaneously recovers both canonical regret scalings known for randomized linear bandit algorithms: the $\tilde O(d^{3/2}\sqrt{T})$ rate for arbitrary bounded arm sets and the $\tilde O (d\sqrt{T\log K})$ rate for finite arm sets. The algorithm also admits an anytime implementation without resetting past data, and experiments show that it remains competitive with baselines while using substantially smaller ensembles.
发表机构
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。