将强化学习扩展至大规模群体:学习平均场表示
Towards Scaling Reinforcement Learning to Massive Populations: Learning Mean-Field Representations
- New York University(纽约大学)
- ETH Zurich(苏黎世联邦理工学院)
- Meta Platforms Inc.(元平台公司)
- Tel Aviv University(特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对大规模群体高维控制问题,提出平均场强化学习框架,通过学习低维群体表示提升奖励预测与策略均衡质量,验证了该方法的有效性。
AI中文摘要:
现代多智能体系统正越来越多地部署在广告拍卖、交通路由和推荐系统等场景的大规模智能体群体中。此类场景的主流方法是独立优化每个智能体的策略,将其他智能体视为固定单智能体环境的一部分,而非对群体动态进行建模。在许多大规模群体系统中,动态取决于群体的聚合摘要,而非任何个体的身份。平均场强化学习(Mean-Field RL)利用这种结构,提供了一个原则性框架,将每个智能体的环境建模为群体分布的显式函数。然而,在状态-动作空间较大或高维控制问题中,对群体分布进行建模本身是难以处理的。本研究从表示学习的角度探讨了“如何为具有大规模群体的高维控制问题设计可扩展框架”这一问题。我们引入了一个平均场强化学习框架,其中奖励和转移动态仅通过未知的低维聚合统计量依赖于群体。随后,我们在离线环境中研究该框架,并设计了一种可证明的方法,通过学习低维表示来学习接近最优的策略。受现实生活中的供应链优化问题启发,我们设计了一个一步路由博弈,以检验“与不利用该结构的基线方法相比,学习低维群体表示可改善奖励预测和纳什间隙估计”这一假设。我们表明,在固定的神经网络参数数量和优化预算下,学习低维群体表示可改善奖励预测以及所得策略的均衡质量。
英文摘要:
Modern multi-agent systems are increasingly deployed at scale over large populations of agents in settings such as ad-auctions, traffic routing, and recommendation systems. The dominant approach in such settings is to optimize each agent's policy independently, treating the other agents as part of a fixed single-agent environment rather than modeling the population dynamics. In many large-population systems, the dynamics depend on an aggregate summary of the population rather than the identity of any individual. Mean-field RL exploits such structure, providing a principled framework that models each agent's environment as an explicit function of the population distribution. However, in large state-action spaces or high-dimensional control problems, modeling the population distribution is itself intractable. How can we design a scalable framework for high-dimensional control problems with large populations? This work explores this question from the perspective of representation learning. We introduce a mean-field RL framework in which the rewards and transition dynamics depend on the population only through an unknown low-dimensional aggregate statistic. We then study this framework in the offline setting and design a provable approach that learns a near-optimal policy by learning a low-dimensional representation. Motivated by real-life supply-chain optimization problems, we design a one-step routing game to test the hypothesis that learning a low-dimensional population representation improves reward prediction and Nash gap estimation relative to baselines that don't exploit this structure. We show that under a fixed neural-network parameter count and optimization budget, learning a low-dimensional population representation improves reward prediction and the equilibrium quality of the resulting policies.