AI 中文总结
研究参数化动作强化学习中多智能体演员-评论家算法,提出共享经验多智能体扩展框架,通过在特定基准上评估,发现多智能体框架提升贪婪演员-评论家性能,多智能体软演员-评论家等提升较小,还揭示了智能体数量增加时性能与计算成本的权衡。
AI 中文摘要
参数化动作强化学习在需要离散动作选择和连续参数化的环境中表现出强大性能。先前工作确立了单智能体演员-评论家算法(贪婪演员-评论家、软演员-评论家、截断分位数评论家)在基准参数化动作任务上的有效性,但它们在多智能体设置中的扩展仍未充分探索。本文对这些算法的共享经验多智能体扩展(多智能体贪婪演员-评论家、多智能体软演员-评论家、多智能体截断分位数评论家)进行比较研究。该框架使用多个共享重放缓冲区但保持独立策略和价值网络的独立演员-评论家智能体。在Platform-v0和Goal-v0基准上针对单智能体对应算法评估这些算法,使用三、五和十个智能体配置评估可扩展性。通过平均评估回报和训练时间衡量性能,用单向方差分析和Tukey HSD事后检验评估统计显著性。结果表明多智能体框架持续提升贪婪演员-评论家性能,而多智能体软演员-评论家及多智能体截断分位数评论家相比单智能体版本提升较小。增加智能体数量超过五个时性能提升有限但计算成本大幅增加,特别是多智能体贪婪演员-评论家。这些结果突出了学习性能和计算效率之间的权衡,为参数化动作强化学习的共享经验多智能体演员-评论家方法的可扩展性提供了见解。
英文摘要
Parameterized action reinforcement learning has shown strong performance in environments requiring both discrete action selection and continuous parameterization. Prior work established the effectiveness of single-agent actor-critic algorithms - Greedy Actor-Critic (GAC), Soft Actor-Critic (SAC), and Truncated Quantile Critics (TQC) - on benchmark parameterized action tasks, but their extension to multi-agent settings remains largely unexplored. This paper presents a comparative study of shared-experience multi-agent extensions of these algorithms: Multi-Agent Greedy Actor-Critic (MAGAC), Multi-Agent Soft Actor-Critic (MASAC), and Multi-Agent Truncated Quantile Critics (MATQC). Rather than following the centralized training, decentralized execution (CTDE) paradigm, the proposed framework uses multiple independent actor-critic agents that share a replay buffer while maintaining separate policy and value networks. We evaluate the algorithms on the Platform-v0 and Goal-v0 benchmarks against their single-agent counterparts, using three-, five-, and ten-agent configurations to assess scalability. Performance is measured by average evaluation return and training time across ten independent runs, with one-way ANOVA and Tukey HSD post-hoc tests used to assess statistical significance. Results show that the multi-agent framework consistently improves Greedy Actor-Critic performance, while MASAC and MATQC show comparatively modest gains over their single-agent versions. Increasing the number of agents beyond five yields limited additional performance while substantially raising computational cost, particularly for MAGAC. These results highlight a trade-off between learning performance and computational efficiency, offering insight into the scalability of shared-experience multi-agent actor-critic methods for parameterized action reinforcement learning.
Comments9 pages, 2 figures