arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

元多智能体强化学习用于交互策略的快速适应及其在自动驾驶中的应用

Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving

Huiwen Yan, Kyriakos G. Vamvoudakis, Mushuang Liu

arXiv 2610.00705首次发表:更新:

发表机构

Virginia Tech; Georgia Tech(弗吉尼亚理工大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出元多智能体强化学习框架,通过马尔可夫博弈建模和元纳什均衡概念,实现自动驾驶中交互策略的快速适应,并验证其优于预训练基线。

AI 中文摘要

本文开发了一个元多智能体强化学习(meta-MARL)框架,以实现多智能体系统(MAS)中交互策略的快速适应。元强化学习(meta-RL)使智能体能够通过双层优化机制快速适应新任务/环境。然而,现有的元强化学习通常侧重于单智能体系统。将这些框架和算法扩展到多智能体系统会带来额外的挑战,因为任务不仅由环境特征化,还由智能体的策略交互特征化。为了应对这些挑战,我们将多智能体强化学习(MARL)问题建模为马尔可夫博弈(MGs),并开发了一个元多智能体强化学习框架,用于在一系列马尔可夫博弈分布中快速适应交互策略。我们定义了一个新概念,称为元纳什均衡(meta-NE),用以描述元多智能体强化学习问题中期望的解概念。我们建立了元纳什均衡与基于梯度博弈的元多智能体强化学习算法的稳定点之间等价性的充分条件。我们在自动驾驶任务上的评估表明,所提出的元多智能体强化学习方法比预训练的多智能体强化学习基线实现了更快的适应,验证了我们框架的有效性。

英文摘要

This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS). Meta-reinforcement learning (meta-RL) enables agents to rapidly adapt to new tasks/environments using a bi-level optimization mechanism. However, existing meta-RL generally focuses on single-agent systems. Extending these frameworks and algorithms to multi-agent systems poses additional challenges, as tasks are characterized by not only the environment but also agents' strategic interactions. To address these challenges, we model multi-agent reinforcement learning (MARL) problems as Markov games (MGs) and develop a meta-MARL framework for rapid interactive policy adaptation across a distribution of MGs. A new concept, called meta-NE, is defined to describe the desired solution concept in a meta-MARL problem. Sufficient conditions for the equivalence between a meta-NE and a stationary point of the gradient-play-based meta-MARL algorithm are established. Our evaluation on autonomous-driving tasks demonstrates that the proposed meta-MARL method achieves faster adaptation than pretrained MARL baselines, validating the effectiveness of our framework.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑