发表机构
Technion – Institute of Technology(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出多臂老虎机中的期望样本复杂度度量,在ACE框架下分析,证明其优于高概率保证,并设计算法及下界,揭示确定性算法劣势与两种设置下的性能分离。
AI 中文摘要
样本复杂度是序贯决策问题中广泛使用的度量,定义为智能体与环境交互过程中次优决策的数量。我们研究随机多臂老虎机问题的样本复杂度,引入期望样本复杂度这一性能度量,并在一个称为期望近似正确(ACE)的新框架中对其进行分析。我们证明,与其他框架中的高概率保证相比,ACE保证蕴含了期望最优奖励的几乎必然收敛,并展示了如何将ACE保证转化为显式的期望遗憾界。我们进一步表明,与现有度量不同,确定性算法无法获得有利的ACE界,并在两种设置下分析随机算法:当允许的次优水平ε对算法已知时和未知时。在前者中,我们设计了一种先探索后ε-贪心算法;在后者中,我们分析了汤普森采样的期望样本复杂度。最后,我们为两种设置建立了近乎匹配的下界,证明这些算法在ε上是紧的,并证明了两种机制之间的性能分离。
英文摘要
Sample complexity is a widely used metric in sequential decision-making problems, defined as the number of suboptimal decisions during the interaction between the agent and an environment. We study the sample complexity of stochastic multi-armed bandit problems and introduce the expected sample complexity performance measure, analyzing it in a novel framework called approximately correct in expectation (ACE). We show that ACE guarantees imply almost sure convergence to the optimal expected reward, in contrast to high-probability guarantees found in other frameworks, and also show how to convert ACE guarantees into explicit expected regret bounds. We further show that, in contrast to existing measures, deterministic algorithms cannot obtain favorable ACE bounds, and analyze stochastic algorithms in two settings: when the allowed suboptimality level $ε$ is known to the algorithm and when it is unknown. In the former, we devise an explore-then-$ε$-greedy algorithm, and in the latter, we analyze the expected sample complexity of Thompson sampling. Finally, we establish nearly matching lower bounds for both settings, showing that the algorithms are tight in $ε$ and proving a performance separation between the two regimes.