arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2301.11518cs.LG

具有全知跟随者的Stackelberg博弈中的在线学习

Online Learning in Stackelberg Games with an Omniscient Follower

  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

Geng Zhao, Banghua Zhu, Jiantao Jiao, Michael I. Jordan

更新

AI总结:

本文研究具有全知跟随者的去中心化合作Stackelberg博弈在线学习问题,揭示了跟随者的全知性可使遗憾最小化的样本复杂度从常数级剧变至指数级,为领导者学习与遗憾分析带来新挑战。

AI中文摘要:

我们研究了一个双玩家去中心化合作Stackelberg博弈中的在线学习问题。在每一轮中,领导者首先采取行动,随后跟随者在观察到领导者的行动后采取自己的行动。领导者的目标是根据交互历史学习以最小化累积遗憾。与传统的重复Stackelberg博弈表述不同,我们假设跟随者是全知的,完全了解真实奖励,并且始终对领导者的行动做出最优响应。我们分析了这一重复Stackelberg博弈中遗憾最小化的样本复杂度。我们证明,取决于奖励结构,全知跟随者的存在可能会极大地改变样本复杂度,从常数级到指数级,即使对于线性合作Stackelberg博弈也是如此。这为领导者的学习过程以及随后的遗憾分析带来了独特的挑战。

英文摘要:

We study the problem of online learning in a two-player decentralized cooperative Stackelberg game. In each round, the leader first takes an action, followed by the follower who takes their action after observing the leader's move. The goal of the leader is to learn to minimize the cumulative regret based on the history of interactions. Differing from the traditional formulation of repeated Stackelberg games, we assume the follower is omniscient, with full knowledge of the true reward, and that they always best-respond to the leader's actions. We analyze the sample complexity of regret minimization in this repeated Stackelberg game. We show that depending on the reward structure, the existence of the omniscient follower may change the sample complexity drastically, from constant to exponential, even for linear cooperative Stackelberg games. This poses unique challenges for the learning process of the leader and the subsequent regret analysis.

↑