多智能体系统的激励设计:协调独立智能体的双层优化框架及收敛性分析
Incentive Design for Multi-Agent Systems: A Bilevel Optimization Framework for Coordinating Independent Agents and Convergence Analysis
- University of Florida(佛罗里达大学)
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对领导者-跟随者多智能体系统,提出基于双层优化的激励设计框架,利用MDP值函数光滑性开发收敛算法,并在随机网格世界中验证其有效性。
AI中文摘要:
激励设计旨在引导系统的性能朝向人类的意图或偏好。我们在一个包含一个领导者和多个跟随者的多智能体系统中研究该问题。每个跟随者独立求解一个马尔可夫决策过程(MDP)以最大化其自身的期望总回报,且所有跟随者具有相同的状态空间和动作空间。然而,领导者的目标取决于所有跟随者的集体最优响应策略。为了影响这些跟随者的策略,领导者以一定成本向各个跟随者提供侧支付作为激励,旨在使跟随者的集体行为与其自身目标对齐,同时最小化该激励成本。这种领导者-跟随者交互被建模为一个双层优化问题:下层由跟随者在给定侧支付的情况下独立优化其MDP组成,上层涉及领导者在给定跟随者最优响应的情况下优化其目标函数。解决激励设计的主要挑战在于领导者的目标通常是非凹的,且下层优化问题可能存在多个局部最优解。为此,我们采用该双层优化问题的约束优化重构形式,并利用MDP中值函数的若干光滑性性质,开发了一种可证明收敛到原始问题驻点的算法。我们在一个随机网格世界中验证了我们的算法,通过检查其收敛性、验证约束得到满足以及评估领导者性能的提升来加以验证。
英文摘要:
Incentive design aims to guide the performance of a system towards a human's intention or preference. We study this problem in a multi-agent system with one leader and multiple followers. Each follower independently solves a mdp to maximize its own expected total return with the same state space and action space. However, the leader's objective depends on the collective best-response policies of all followers. To influence these policies of followers, the leader provides side payments as incentives to individual followers at a cost, aiming to align the collective behaviors of followers with its own goal while minimizing this cost of incentive. Such a leader-followers interaction is formulated as a bilevel optimization problem: the lower level consists of followers individually optimizing their MDPs given the side payments, and the upper level involves the leader optimizing its objective function given the followers' best responses. The main challenge to solve the incentive design is that the leader's objective is generally non-concave and the lower level optimization problems can have multiple local optima. To this end, we employ a constrained optimization reformation of this bi-level optimization problem and develop an algorithm that provably converges to a stationary point of the original problem, by leveraging several smoothness properties of value functions in MDPs. We validate our algorithm in a stochastic gridworld by examining its convergence, verifying that the constraints are satisfied, and evaluating the improvement in the leader's performance.