PLATO:用于智能体和任务开放性的指针学习者
PLATO: Pointer Learner for Agent and Task Openness
浏览论文内容
中文总结 AI 辅助
研究开放智能体系统中智能体和任务开放性带来的多智能体强化学习挑战,提出基于指针网络的PLATO方法,结合集中式GNN评论家,在无界状态和动作空间下训练,在野火抑制领域评估中性能优于现有基线。
中文摘要 AI 辅助
开放智能体系统(OASYS)在现实世界领域日益普遍,其中智能体和任务集随时间变化不可预测。这种开放性对多智能体强化学习(MARL)构成挑战,现有方法仅部分解决开放性问题。本文介绍了用于智能体和任务开放性的指针学习者(PLATO),它是基于指针网络的智能体与集中式图神经网络(GNN)评论家相结合,在集中训练和分散执行范式下用多智能体近端策略优化进行训练。基于指针的智能体直接输出当前任务集上的分布,GNN评论家将智能体 - 任务交互编码为随任务和智能体组成变化形状的图。该方法在无界状态和动作空间中考虑了智能体开放性和任务开放性,通过野火抑制领域评估展示了优于现有基线的性能和零样本泛化能力。
英文摘要
Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to multi-agent reinforcement learning (MARL), which typically assumes fixed state and action spaces. Existing methods address openness only partially: padding and masking approaches introduce artificial bounds, while recent graph-based or hypergraph methods handle one dimension of openness but still depend on restrictive assumptions. In this paper, we introduce Pointer Learner for Agent and Task Openness (PLATO), a pointer-network-based actor combined with a centralized graph neural network (GNN) critic, trained with multi-agent proximal policy optimization under a centralized training and decentralized execution paradigm. Our pointer-based actor outputs distributions directly over the current task set. This directly supports changing action spaces without masking or retraining. Our GNN critic encodes agent-task interactions as a graph that changes shape with task and agent composition. Together, these components consider AO and TO without the boundedness of existing approaches. We formalize PLATO in a Task-and-Agent-Open Markov Game (TaAgO-MG), extending prior task-open formulations, and prove it is well-defined over the resulting unbounded state and action spaces. We evaluate PLATO with the Methods for Open Agent Systems Evaluation Initiative (MOASEI) wildfire suppression domain, an environment designed for open multi-agent system evaluation, and we demonstrate strong performance and more consistent zero-shot generalization than state-of-the-art baselines in OASYS.
发表机构
- University of Nebraska-Lincoln(内布拉斯加大学林肯分校)
- University of Georgia(佐治亚大学)
- Oberlin College(欧柏林学院)
机构由 AI 辅助整理,请以论文原文为准。