在未知追随者类型下的贝叶斯斯塔克尔伯格博弈中的学习
Learning in Bayesian Stackelberg Games With Unknown Follower's Types
- Politecnico di Milano(米兰理工大学)
- MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究了在未知追随者类型情况下,设计无遗憾算法以最小化领导者遗憾的贝叶斯斯塔克尔伯格博弈学习问题。
AI中文摘要:
我们研究了贝叶斯斯塔克尔伯格博弈中的在线学习,其中领导者反复与一个未知私人类型的追随者互动,该类型在每个回合中独立地从未知的概率分布中抽取。目标是设计出在知道游戏的情况下始终玩最优承诺的算法,以最小化领导者的遗憾。我们首次考虑了最现实的情况,即领导者对追随者的类型一无所知,即可能的追随者支付情况。这比通常研究的情况更具挑战性,即追随者类型的支付情况是已知的。首先,我们证明了一个强负结果:在仅观察追随者在每个回合结束时的最佳反应的情况下,无遗憾是无法实现的。因此,我们专注于更容易的类型反馈模型,其中追随者的类型也被揭示。在这种情况下,我们提出了一种无遗憾算法,其遗憾为O(√T),当忽略对其他参数的依赖时。
英文摘要:
We study online learning in Bayesian Stackelberg games, where a leader repeatedly interacts with a follower whose unknown private type is independently drawn at each round from an unknown probability distribution. The goal is to design algorithms that minimize the leader's regret with respect to always playing an optimal commitment computed with knowledge of the game. We consider, for the first time to the best of our knowledge, the most realistic case in which the leader does not know anything about the follower's types, i.e., the possible follower payoffs. This raises considerable additional challenges compared to the commonly studied case in which the payoffs of follower types are known. First, we prove a strong negative result: no-regret is unattainable under action feedback, i.e., when the leader only observes the follower's best response at the end of each round. Thus, we focus on the easier type feedback model, where the follower's type is also revealed. In such a setting, we propose a no-regret algorithm that achieves a regret of $\widetilde{O}(\sqrt{T})$, when ignoring the dependence on other parameters.