谁承担负担?多智能体强化学习中共享约束的责任学习
Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对多智能体共享成本约束的惩罚分配问题,提出LiRA方法,通过学习共同乘子的责任份额优化社会福利,在多个基准上提升平均福利达29%。
中文摘要 AI 辅助
当多个智能体共享一个成本预算时,一个共同的拉格朗日乘子可以强制执行聚合约束,但并未决定其惩罚应如何在智能体之间分配。均匀惩罚忽略了智能体所牺牲奖励的异质性,而智能体特定的乘子可能仍然依赖于相同的聚合成本信号。我们引入了拉格朗日责任分配(LiRA),它通过在一个有限训练范围内优化社会福利来学习每个智能体对共同乘子的份额。乘子强制执行聚合预算,而责任份额重新分配其影响,而不修改原始奖励或约束。对于标准正则条件下的凸博弈,改变这些份额会产生一个平滑的归一化广义纳什均衡族,其中活跃约束保持在预算内,而福利发生变化。为了在收敛前优化责任,我们推导了一个福利梯度,该梯度同时考虑了学习更新和引起的数据分布变化。在CityLearn、MABIM、Harvest和MetaDrive中,涵盖3到400个智能体,LiRA相比均匀和智能体特定的乘子基线,平均社会福利提高了高达29%。网格和驾驶成本保持在预算内,库存违规减少,Harvest更有效地利用了可用预算。
英文摘要
When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multipliers may still rely on the same aggregate cost signal. We introduce Lagrangian Responsibility Allocation (LiRA), which learns each agent's share of a common multiplier by optimizing social welfare over a finite training horizon. The multiplier enforces the aggregate budget, while responsibility shares redistribute its influence without modifying the original rewards or constraints. For convex games under standard regularity conditions, varying these shares induces a smooth family of normalized generalized Nash equilibria in which active constraints remain at their budgets while welfare varies. To optimize responsibility before convergence, we derive a welfare gradient that accounts for both learning updates and the induced change in data distribution. Across CityLearn, MABIM, Harvest, and MetaDrive, spanning 3 to 400 agents, LiRA improves average social welfare by up to 29% over uniform and agent-specific multiplier baselines. Grid and driving costs remain within budget, inventory violations decrease, and Harvest makes more effective use of available budget.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- Stanford University(斯坦福大学)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。