带辅助信息的Stackelberg博弈中的遗憾最小化
Regret Minimization in Stackelberg Games with Side Information
- School of Computer Science(计算机学院)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出带辅助信息的Stackelberg博弈,证明完全对抗场景下领导者无法实现无遗憾,但在两种混合随机与对抗性的放松场景中可实现无遗憾学习。
AI中文摘要:
用于参与Stackelberg博弈的算法已被部署在包括机场安全、反偷猎行动和网络犯罪预防等现实领域的应用中。然而,这些算法往往未能考虑到每个玩家可获取的额外信息(例如交通模式、天气状况、网络拥堵),这些信息可能会显著影响双方玩家的最优策略。我们将此类场景形式化为带辅助信息的Stackelberg博弈,其中双方玩家在博弈前都会观察到一个外部上下文。领导者承诺一种(依赖于上下文的)策略,而跟随者则对领导者的策略和上下文做出最优响应。我们关注在线场景,即一系列跟随者随时间到达,且上下文可能在每一轮之间发生变化。与非上下文版本形成鲜明对比的是,我们表明在完全对抗性场景中,领导者不可能实现无遗憾。受此结果启发,我们证明了在两种自然放松的场景下无遗憾学习是可行的:一种是跟随者序列被随机选择且上下文序列是对抗性的场景,另一种是上下文是随机的且跟随者类型是对抗性的场景。
英文摘要:
Algorithms for playing in Stackelberg games have been deployed in real-world domains including airport security, anti-poaching efforts, and cyber-crime prevention. However, these algorithms often fail to take into consideration the additional information available to each player (e.g. traffic patterns, weather conditions, network congestion), which may significantly affect both players' optimal strategies. We formalize such settings as Stackelberg games with side information, in which both players observe an external context before playing. The leader commits to a (context-dependent) strategy, and the follower best-responds to both the leader's strategy and the context. We focus on the online setting in which a sequence of followers arrive over time, and the context may change from round-to-round. In sharp contrast to the non-contextual version, we show that it is impossible for the leader to achieve no-regret in the full adversarial setting. Motivated by this result, we show that no-regret learning is possible in two natural relaxations: the setting in which the sequence of followers is chosen stochastically and the sequence of contexts is adversarial, and the setting in which contexts are stochastic and follower types are adversarial.