AI 中文总结
本文针对开放多智能体系统分布式任务分配问题,提出SubMAPL方法,通过划分多线性扩展解决次模函数松弛与分类策略不匹配问题,在多智能体覆盖任务上性能优于基线方法。
AI 中文摘要
本文研究开放多智能体系统中分布式任务分配的策略学习问题,其中智能体以时变方式加入和离开,阶段团队效用为次模函数。在每个时间步,活跃智能体从局部分类策略中选择动作,使得可行的联合智能体-动作对构成一个划分拟阵。次模集函数的标准连续松弛基于独立伯努利采样,这与智能体的动作选择不匹配。为解决该不匹配问题,我们提出划分多线性扩展(Partition Multilinear Extension, PME),这是一种基于策略的松弛,其连续支持与分类策略下的可行动作匹配。我们证明阶段效用的边际收益是PME梯度的无偏估计,且在动作分布上最大化PME等价于在智能体动作上最大化阶段效用,这对设计有原则的策略至关重要。基于此,我们设计SubMAPL,这是一种集中训练、分散执行的KL镜像策略学习方法,在训练期间使用局部边际收益作为随机PME梯度。KL镜像更新保持分类可行性,无需欧几里得投影。在智能体运行表格型-softmax策略的情况下,我们引入开放策略迁移和开放系统KL跟踪变体以处理智能体的加入和离开。通过动态后悔分析,我们建立了累积效用的下界,该下界考虑了环境的开放性以及最优阶段效用与全局效用之间的差距。在多智能体覆盖问题上的仿真表明,SubMAPL的性能优于策略梯度和在线学习基线方法。
英文摘要
This paper studies policy learning for distributed task allocation in open multi-agent systems, where agents may join and leave in a time-varying fashion, with submodular stage team utilities. At each time, the active agents select actions from local categorical policies such that the feasible joint agent-action pairs form a partition matroid. Standard continuous relaxations of submodular set functions are based on independent Bernoulli sampling, making them inconsistent with agents' policies.To solve this mismatch, we propose the \emph{partition multilinear extension} (PME), a policy-based relaxation whose continuous support matches feasible actions under categorical policies.We prove that the marginal gains of the stage utility provide an unbiased estimator of the gradient of the PME and that maximizing the PME over action distributions is equivalent to maximizing the stage utilities over agent actions, which are critical to devise principled policy gradient.Building on this, we design \emph{SubMAPL}, a centralized-training decentralized-execution KL-mirror policy-learning method that uses local marginal gains as stochastic PME gradients during training. KL-mirror updates preserve categorical feasibility without Euclidean projection.In the case where agents run tabular-softmax policies, we introduce open policy migration and an open-system KL tracking variation to handle agent arrivals and departures. Using dynamic regret analysis, we establish a lower bound on the cumulative utility which accounts for the openness of the environment and for the gap between optimal stage-wise and global utilities. Simulations on multi-agent coverage demonstrate that SubMAPL outperforms policy-gradient and online-learning baselines.