学习多跟随者贝叶斯斯塔克尔伯格博弈
Learning to Play Multi-Follower Bayesian Stackelberg Games
- John A. Paulson School of Engineering and Applied Sciences(约翰·A·保罗森工程与应用科学学院)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究多跟随者贝叶斯斯塔克尔伯格博弈的在线学习算法,通过类型反馈和动作反馈设计算法,实现最小化懊悔的策略。
AI中文摘要:
在多跟随者贝叶斯斯塔克尔伯格博弈中,领导者在$L$个动作上采用混合策略,而每个跟随者具有$K$种可能的私人类型之一,并对其最佳反应。领导者最优策略取决于跟随者私人类型的分布。我们研究该问题的在线学习版本:领导者与$n$个跟随者进行$T$轮互动,每个轮次中跟随者的类型均从未知分布中采样。领导者的目标是最小化懊悔,即最优策略累积效用与实际选择策略累积效用的差值。我们为领导者在不同反馈设置下设计了学习算法。在类型反馈设置下,领导者在每轮结束后观察跟随者的类型,我们设计了在独立类型分布下达到$O(\sqrt{\min(L\log(nKA T), nK ) \cdot T})$懊悔,在一般类型分布下达到$O(\sqrt{\min(L\log(nKA T), K^n ) \cdot T})$懊悔的算法。有趣的是,这些界不随$n$以多项式速率增长。在动作反馈设置下,领导者仅观察跟随者的行为,我们设计了达到$O( \min(\sqrt{ n^L K^L A^{2L} L T \log T}, K^n\sqrt{ T } \log T ) )$懊悔的算法。我们还提供了下界$Ω(\sqrt{\min(L, nK)T})$,几乎匹配类型反馈上界。
英文摘要:
In a multi-follower Bayesian Stackelberg game, a leader plays a mixed strategy over $L$ actions to which $n\ge 1$ followers, each having one of $K$ possible private types, best respond. The leader's optimal strategy depends on the distribution of the followers' private types. We study an online learning version of this problem: a leader interacts for $T$ rounds with $n$ followers with types sampled from an unknown distribution every round. The leader's goal is to minimize regret, defined as the difference between the cumulative utility of the optimal strategy and that of the actually chosen strategies. We design learning algorithms for the leader under different feedback settings. Under type feedback, where the leader observes the followers' types after each round, we design algorithms that achieve $O\big(\sqrt{\min(L\log(nKA T), nK ) \cdot T} \big)$ regret for independent type distributions and $O\big(\sqrt{\min(L\log(nKA T), K^n ) \cdot T} \big)$ regret for general type distributions. Interestingly, those bounds do not grow with $n$ at a polynomial rate. Under action feedback, where the leader only observes the followers' actions, we design algorithms with $O( \min(\sqrt{ n^L K^L A^{2L} L T \log T}, K^n\sqrt{ T } \log T ) )$ regret. We also provide a lower bound of $Ω(\sqrt{\min(L, nK)T})$, almost matching the type-feedback upper bounds.