合作多智能体强化学习的在线变点检测
Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning
- University of Saskatchewan(萨斯喀彻温大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对合作多智能体强化学习提出基于奖励的轻量变点检测算法PPR,经实验验证其能平衡检测速度与警报稳定性,助力MARL系统识别训练中的重大变化。
AI中文摘要:
合作多智能体强化学习(MARL)系统依赖过往经验学习协同行为,但在训练过程中若环境或任务目标发生变化,这些经验可能变得不可靠。在这种情况下,智能体首先需要一种方法来识别情况已发生变化,之后再决定如何适应。本文研究基于奖励衍生信号的合作MARL在线变点检测,提出了“过往奖励模式”(PPR),这是一种轻量、与算法无关的检测器,它对智能体的回报流进行平滑处理,突出近期变化,并应用统计漂移检测器标记显著转变。我们在基于多智能体粒子环境(Multi-Agent Particle Environment)构建的自定义“说话者-倾听者”环境中,在两种受控非平稳场景下评估PPR。结果显示检测速度与警报稳定性之间存在权衡:平滑回报基线检测更早,但会产生许多重复警报;相反,直接对原始回报应用检测器往往会遗漏转变。PPR通过限制冗余检测同时仍能识别受控转变,提供了一种更平衡的方法。这些发现表明PPR是一种轻量、基于奖励的监控工具,可使合作MARL系统在训练过程中可靠识别重大变化。
英文摘要:
Cooperative multi-agent reinforcement learning (MARL) systems rely on past experience for learning coordinated behaviour, but this experience may become unreliable if the environment or task objective changes during training. In such cases, agents first need a way to recognize that the situation has changed before deciding how to adapt. This paper studies online change-point detection for cooperative MARL using reward-derived signals. We propose \emph{Patterns of Past Rewards} (PPR), a lightweight algorithm-agnostic detector that smooths agents' return streams, highlights recent changes, and applies a statistical drift detector to flag significant shifts. We evaluate PPR in a custom Speaker-Listener environment based on the Multi-Agent Particle Environment under two controlled non-stationarity scenarios. Our results show a trade-off between detection speed and alarm stability. A smoothed-return baseline detects earlier but produces many repeated alarms. In contrast, applying the detector directly to raw returns often misses the shift. PPR offers a more balanced approach by limiting redundant detections while still identifying the controlled shifts. These findings highlight PPR as a lightweight, reward-based monitoring tool that enables cooperative MARL systems to reliably identify major changes during training.