改进的伯努利奖励多人赌博机算法
Improved Multiplayer Bandit Algorithm for Bernoulli Rewards
AI总结:
针对伯努利奖励下多人多臂赌博机信息不对称问题,提出基于KL散度的mKL-UCB等算法,获得更紧的遗憾界,改进因子至少为2。
AI中文摘要:
我们研究了伯努利奖励下具有信息不对称性的多人多臂赌博机问题,针对三种信息结构:动作不对称、奖励不对称以及两者均不对称。用基于Kullback-Leibler(KL)散度的置信区间替代先前工作中Hoeffding风格的置信区间,在每种情况下都严格地获得了更紧的遗憾界。我们提出了mKL-UCB、mKL-UCB-Intervals和mKL-DSEE,并通过Pinsker不等式证明改进因子至少为2,当奖励均值接近0或1时改进幅度更大。对于奖励不对称情况,我们证明两个臂的KL区间在确定数量的样本后分离,且M个独立玩家进一步加速消除过程。
英文摘要:
We study the multiplayer multi-armed bandit problem with information asymmetry under Bernoulli rewards, for three information structures: asymmetry in actions, in rewards, and in both. Replacing the Hoeffding-style confidence intervals of prior work with Kullback--Leibler (KL) divergence-based bounds gives strictly tighter regret guarantees in each case. We propose \texttt{mKL-UCB}, \texttt{mKL-UCB-Intervals} and \texttt{mKL-DSEE}, and show that the improvement factor is at least two by Pinsker's inequality and far larger when reward means are near zero or one. For asymmetry in rewards we prove that two arms' KL intervals separate after a deterministic number of samples, and that $M$ independent players accelerate elimination further.