具有连续动作空间的合作任务的自进化默认动作
A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space
浏览论文内容
中文总结 AI 辅助
研究连续动作空间合作任务中反事实信用分配难题,提出SAFE框架,利用自进化默认动作构建反事实基线,解决偏差和收敛问题,实验证明其性能优于现有模型。
中文摘要 AI 辅助
反事实信用分配在离散动作空间的多智能体强化学习(MARL)中已被证明是有效的,但其扩展到连续动作合作任务仍然具有挑战性。现有通过蒙特卡罗采样近似反事实基线的方法,由于采样动作可能未得到充分训练,常将偏差引入策略梯度且无法保证收敛到局部最优。为解决这些限制,我们提出了SAFE,这是一个新颖的MARL框架,它采用基于从每个智能体经验缓冲区采样的自进化默认动作的反事实基线。该设计自然地扩展到连续动作空间,无需依赖额外模拟、奖励模型或特定于环境的先验知识。基线准确量化每个智能体的贡献,且不会将偏差引入确定性政策梯度,确保收敛到局部最优。在合作车辆任务上的大量实验表明,SAFE始终优于现有模型。
英文摘要
Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that approximate the counterfactual baseline via Monte Carlo sampling often introduce bias into policy gradients and fail to guarantee convergence to local optima, as the sampled actions may not have been sufficiently trained. To address these limitations, we propose SAFE, a novel MARL framework that employs a counterfactual baseline conditioned on a self-evolving default action sampled from each agent's experience buffer. This design naturally extends to continuous action spaces without relying on additional simulations, reward models, or environment-specific prior knowledge. The baseline accurately quantifies each agent's contribution, and introduces no bias into the deterministic policy gradient, ensuring convergence to local optima. Extensive experiments on cooperative vehicular tasks demonstrate that SAFE consistently outperforms state-of-the-art models.
发表机构
- XJTLU(西交利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。