发表机构
Adobe Research(奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种确定性和非耦合的学习动力学,在一般和博弈中实现常数个体交换遗憾,并通过预测偏离收益和更新转移矩阵,同时提供对抗鲁棒变体。
AI 中文摘要
我们给出了在完全信息反馈下,有限多人一般和博弈的确定性和非耦合学习动力学,实现了常数个体交换遗憾,与时间范围 $T$ 无关。对于 $n$ 个玩家且每个玩家至多有 $m$ 个动作的情况,每个玩家在任意有限时间范围内的个体交换遗憾为 $O(\sqrt{n} m \log m \log^{5/2}(nm))$。每个玩家预测偏离收益,然后利用这些预测更新一个行随机转移矩阵,并按其平稳分布进行博弈。证明结合了利用平稳性的势参数与双尺度高阶预测分析,使用有根树表示来处理偏离收益对平稳分布的非线性依赖。通过一个通用的公共前缀切换包装器获得的对抗鲁棒变体,将自博弈界保持到通用常数,并保证在对抗设置中个体交换遗憾至多为 $7\sqrt{m T \log m}$。
英文摘要
We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret in self-play, independent of the horizon $T$. With $n$ players and at most $m$ actions each, every player's individual swap regret is $O(\sqrt n\,m\log m\log^{5/2}(nm))$ at every finite horizon. The dynamics use the classical Blum-Mansour framework with optimism. Each player predicts the deviation gains, uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. Our new ingredients include a tailored row normalization map and a two-scale higher-order predictor. An adversarially robust variant, obtained through a generic common-prefix switching wrapper, preserves the self-play bound up to a universal constant and guarantees individual swap regret at most $7\sqrt{mT\log m}$ in the adversarial setting.