MA-FPPO:多智能体流预训练策略优化
MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization
浏览论文内容
中文总结 AI 辅助
提出MA-FPPO,通过在线微调流匹配预训练模型,利用共享团队优势提升多智能体在离线数据外情境中的合作与协调,平均相对提升52.8%优于离线基线,29.8%优于纯在线学习。
中文摘要 AI 辅助
多智能体流策略从固定的离线数据集中学习合作行为,但往往难以完成离线数据未覆盖情境下的任务。在这些情境中,智能体必须既适应环境变化又相互协调,然而离线学习到的动作模式往往不足以实现有效的适应与协调。为解决这一问题,我们提出了多智能体流预训练策略优化(MA-FPPO),该方法通过在线微调,利用与环境的新交互来改进经流匹配预训练模型的合作行为。基于预训练期间学到的行为,我们为离散和连续动作空间构建了具有显式动作似然的策略。随后,我们使用共享的团队优势来更新预训练模型,以基于团队表现进一步提升协调性。我们的方法在30种设置下,相较于最强的列出的离线基线,平均相对提升达52.8%;在与匹配的在线预算和评估协议进行的38项比较中,相较于纯在线学习,平均相对提升达29.8%。
英文摘要
Multi-agent flow policies learn cooperative behavior from fixed offline datasets, but often struggle to complete tasks in situations not covered by the offline data. In these situations, agents must both adapt to changes in the environment and coordinate with one another, yet action patterns learned offline are often insufficient for effective adaptation and coordination. To address this problem, we propose Multi-Agent Flow-Pretrained Policy Optimization (MA-FPPO), which uses online fine-tuning to improve the cooperative behavior of models pretrained with flow matching through new interactions with the environment. Building on the behavior learned during pretraining, we construct policies with explicit action likelihoods for discrete and continuous action spaces. We then update the pretrained model using shared team advantages to further improve coordination based on team performance. Our method achieves, on average, relative gains of 52.8% over the strongest listed offline baselines across 30 settings and 29.8% over purely online learning across 38 comparisons with matched online budgets and evaluation protocols.