arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01951cs.LGstat.ML

β-EB-TCI 中带惩罚挑战者的尖锐非渐近分析(针对伯努利臂)

Sharp Non-Asymptotic Analysis of the Penalized Challenger in $β$-EB-TCI for Bernoulli Bandits

Nam Nguyen, Tuan Quang Dam

首次发表
浏览论文内容

中文总结 AI 辅助

本文对伯努利臂的 β-EB-TCI 算法进行尖锐非渐近分析,证明其停止时间上界并给出期望样本复杂度,同时指出处理相等次优均值证明策略的局限。

中文摘要 AI 辅助

Top-two 算法对于固定置信度最优臂识别而言简单有效,但其尖锐的非渐近行为仍未被充分理解。我们通过 β-EB-TCI(Jourdan 等人的经验最优 top-two 规则)研究伯努利臂的这一问题的尖锐非渐近行为。该规则中的挑战者通过带有对数计数惩罚的伯努利运输成本来选择。我们证明,在经验领导者已成为真正最优臂且其采样比例保持在 β 附近之后,停止时间等于 T_β^*(μ) log(1/δ),并加上低阶集中项。我们还表明,在此状态下,每个挑战者都被线性频率地采样。因此,对于无强制探索的原始算法,剩余的主要困难在于控制经验领导者何时永久正确。这些结果对所有具有唯一最优臂的伯努利实例给出了非渐近的高概率界。若算法满足有限均值充分探索条件,该界进一步可导出尖锐的期望样本复杂度。特别地,当所有臂均值两两不同时,利用 Jourdan 等人的充分探索结果,这为无防护的伯努利规则提供了尖锐的期望结果。最后,若我们加入一个温和的强制探索规则(该规则在时间 t 前仅贡献 O(√(Kt)) 次拉取),则在唯一最优臂假设下,我们为任意数量的臂获得一个自包含的期望样本复杂度定理。我们还指出了试图通过单一索引比较参数来处理相等次优均值的证明策略的局限性。

英文摘要

Top-two algorithms are simple and effective for fixed-confidence best-arm identification, but their sharp non-asymptotic behavior is still not well understood. We study this problem for Bernoulli bandits through $β$-EB-TCI, the empirical-best top-two rule of Jourdan et al., whose challenger is chosen using a Bernoulli transportation cost with a logarithmic count penalty. We prove that, after the empirical leader has become the true best arm and its sampling fraction stays close to $β$, the stopping time is $T_β^{\star}(μ)\log(1/δ)$ up to lower-order concentration terms. We also show that, in this regime, every challenger is sampled linearly often. Thus, for the original algorithm without forced exploration, the main remaining difficulty is to control when the empirical leader becomes permanently correct. These results imply a non-asymptotic high-probability bound for all Bernoulli instances with a unique best arm. If the algorithm satisfies a finite-mean sufficient-exploration condition, the bound further yields the sharp expected sample complexity. In particular, this gives the sharp expectation result for the unguarded Bernoulli rule when all arm means are pairwise distinct, using the sufficient-exploration result of Jourdan et al. Finally, if we add a mild forced-exploration rule that contributes only $O(\sqrt{Kt})$ pulls up to time $t$, we obtain a self-contained expected sample-complexity theorem for any number of arms under the unique-best-arm assumption. We also identify a limitation of proof strategies that try to handle equal suboptimal means through a single index-comparison argument.

发表机构

  • Hanoi University of Science and Technology(河内理工大学)
  • Center for AI Research, VinUniversity(文大人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑