arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12389stat.MLcs.LG

汤普森采样在错误率上具有2倍竞争力

Thompson Sampling Is 2-Competitive for Mistakes

Mark Sellke, Gregory Valiant

首次发表
浏览论文内容

中文总结 AI 辅助

研究贝叶斯博弈模型,证明汤普森采样在错误率方面最多是其他策略的两倍,适用于特定条件下的随机博弈,证实相关猜想,且结果在多种权重序列下成立。

中文摘要 AI 辅助

我们考虑贝叶斯博弈模型,并证明汤普森采样所犯错误(选择次优臂)的预期数量最多是任何其他策略的两倍。只要潜在臂过程是独立的,且每个臂仅在被选择时才演变,我们的分析就适用。对于通过平均奖励定义最佳臂的随机博弈,这证实了古哈和穆纳加拉2014年的一个猜想,其中2这个因子已经是最优的。该结果在任何非递增的轮次权重序列下都成立,包括固定 horizon 和几何折扣。

英文摘要

We consider Bayesian bandit models and prove that Thompson sampling makes at most twice the expected number of mistakes (selections of a suboptimal arm) as any other policy. Our analysis applies as long as the latent arm processes are independent and each arm evolves only when played. For stochastic bandits with best arm defined via mean reward, this confirms a conjecture of Guha and Munagala from 2014, where the factor $2$ is already best possible. The result holds under any nonincreasing sequence of round weights, including fixed horizon and geometric discounting.

补充信息

↑