发表机构
Univ. Grenoble Alpes; CNRS; Inria; Grenoble INP; LIG(格勒诺布尔阿尔卑斯大学; 法国国家科学研究中心; 法国国家信息与自动化研究所; 格勒诺布尔国立理工学院; 格勒诺布尔信息学实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文围绕博弈学习的遗憾、均衡与学习,结合单智能体和多智能体场景,研究正则化学习策略及相关理论,为该领域提供入门性阐述。
AI 中文摘要
本笔记旨在作为博弈学习相关文献的入门,该主题兼具重要的理论吸引力与广泛的应用场景,覆盖机器学习、数据科学乃至经济学等诸多领域。我们的阐述围绕两个互补视角展开:首先,我们考虑单个智能体——学习者——在未知、非平稳且可能为对抗性的环境中参与序贯决策过程;随后,我们探究当环境由多个相互作用的智能体的决策塑造时会发生什么,这些智能体未必知晓彼此的行动或目标,且均致力于提升自身的个体收益。在这一通用语境下,我们研究一类正则化学习策略,该策略基于对过往博弈历史的最优响应,同时引入正则化惩罚项以鼓励探索、避免对次优选择的过度依赖。在单智能体场景中,我们给出对抗性多臂老虎机中正则化学习的基础遗憾界;在多智能体场景中,我们描述零和博弈的遍历均衡收敛结果,其符合虚拟博弈的经典结论,同时阐述关联策略与动态稳定性概念的“民间定理”——分别对应纳什均衡与正则化学习的吸引点。我们特别关注智能体可获取的信息,并通过统一分析框架研究基于神谕与基于收益(老虎机)的两类方法。我们的目标是为该领域的部分最新思想提供连贯且易懂的(尽管必然不全面的)阐述,并讨论其对理性研究的意义。
英文摘要
This note aims to serve as an entry point to the literature on learning in games, a topic with significant theoretical appeal and a wide range of applications -- from machine learning and data science to economics and beyond. Our presentation is structured around two complementary viewpoints: We first consider a single agent -- the learner -- engaged in a sequential decision process in an unknown, non-stationary, and possibly adversarial environment. We then examine what happens when the environment is shaped by the decisions of several interacting agents, not necessarily aware of each other's actions or goals, and all seeking to improve their individual rewards. In this general context, we examine a family of regularized learning policies based on best-responding to the past history of play, up to a regularization penalty intended to encourage exploration and prevent over-commitment to suboptimal choices. In the single-agent setting, we present some basic regret bounds for regularized learning in adversarial multi-armed bandits; in the multi-agent setting, we describe an ergodic equilibrium convergence result for zero-sum games in the spirit of classical results on fictitious play, as well as a "folk theorem" linking strategic and dynamic notions of stability -- Nash equilibria and attracting points of regularized learning, respectively. We pay special attention to the information available to the players and, through a unified analysis framework, we study both oracle- and payoff-based (bandit) methods. Our goal is to provide a coherent and comprehensible -- albeit, by necessity, not comprehensive -- account of some recent ideas in the field, and to discuss their implications for the study of rationality.
Comments44 pages, 3 figures; to appear as a chapter in "Equilibria in Games: Existence, Selection, and Dynamics", edited by Sylvain Sorin and Bernhard von Stengel