发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种无碰撞信息、无共享随机性的对抗性多人赌博机协议,通过蒙特卡洛公共构造函数实现同步,达到平方根遗憾界。
AI 中文摘要
我们研究了具有 $K$ 个臂和 $2\le m<K$ 个标记玩家的对抗性多人赌博机问题,该问题没有碰撞信息、共享随机性或外部通信渠道。我们设计了一个建设性的通信和同步协议,并采用蒙特卡洛公共构造函数。在预处理阶段,以至少 $1-CN^{-32}$ 的概率(其中 $N=2Km(T+1)$),其固定的公开输出同时满足 \\[ R_T\le C K^{5/2}\sqrt T\log^2(2Km(T+1)) \\] 对于预处理后选择的每个 oblivious 奖励序列。这里 $R_T$ 是玩家私有执行随机性上的期望遗憾。正向奖励观测建立了一个共同的学习时间表,并在学习开始前同步玩家。延迟通信的成本被计入正向奖励的支持,确保反馈较少的时期仅产生有限的遗憾。随后,一个慢-快学习过程在交换分配和分数时维持有效的奖励估计。
英文摘要
We study adversarial multiplayer bandits with $K$ arms and $2\le m<K$ labeled players, without collision information, shared randomness, or an external communication channel. We design a constructive communication and synchronization protocol with a Monte Carlo public constructor. With probability at least $1-CN^{-32}$ over preprocessing, where $N=2Km(T+1)$, its fixed published output satisfies \[ R_T\le C K^{5/2}\sqrt T\log^2(2Km(T+1)) \] simultaneously for every oblivious reward sequence chosen after preprocessing. Here $R_T$ is expected regret over the players' private execution randomness. Positive reward observations establish a common learning schedule and synchronize players before learning begins. The cost of delayed communication is charged to the support of positive rewards, ensuring that periods with little useful feedback incur only limited regret. A slow--fast learning procedure then maintains valid reward estimates while assignments and scores are exchanged.
Comments84 pages, 2 figures