发表机构
Istanbul Technical University; University of Stuttgart; Istanbul Technical University Artificial Intelligence and Data Science Application and Research Center; Stanford University(伊斯坦布尔技术大学; 斯图加特大学; 伊斯坦布尔技术大学人工智能与数据科学应用与研究中心; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通信丢失时多智能体协调问题,提出价值感知MARO方法,利用优势估计动态加权预测器损失函数,在多智能体粒子环境实验中,该方法在通信可靠性下降时能维持性能,提升平均回报并降低性能方差。
AI 中文摘要
鲁棒多智能体协调严重依赖智能体间通信,而在实际部署中常受物理和环境约束干扰。为在通信间歇性故障时维持运行,智能体可采用内部预测模型估计缺失的共享状态信息。但用标准重建目标训练的预测器对所有转移一视同仁。本文提出通信丢包下多智能体观察共享(MARO)的价值感知扩展方法,即价值感知MARO。通过利用底层演员-评论家架构的优势估计动态加权预测器的损失函数,使预测器学习过程与策略演变明确耦合。在多智能体粒子环境的多个任务上评估该框架,实验结果表明在通信可靠性下降时,尤其是低于40%时,该方法能维持性能。在高损耗场景下,与标准未加权基线相比,平均回报提高超20%,性能方差平均降低64.7%。
英文摘要
Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments. To maintain operation during these intermittent communication failures, agents can employ internal prediction models to estimate missing shared state information. However, predictors trained with standard reconstruction objectives treat all transitions equally. In a Reinforcement Learning context, this forces the model to waste capacity learning stochastic exploration noise and the outdated dynamics of suboptimal policies. In this paper, we propose a value-aware extension of Multi-Agent Observation Sharing under Communication Dropout (MARO) to patch communication gaps; we refer to this method as Value-Aware MARO. By dynamically weighting the predictor's loss function using advantage estimates derived from the underlying actor-critic architecture, our objective explicitly couples the predictor's learning process to the policy's evolution. This formulation focuses the model's capacity on the intentional, high-return dynamics actively reinforced by the agents. We evaluate our framework on several tasks within the Multi-Agent Particle Environment under varying communication reliability levels. Experimental results demonstrate that our approach maintains performance under declining communication reliability, particularly below 40%. While our method performs comparably in tasks where the baseline already maintains high coordination, our value-aware weighting effectively prevents the performance collapse observed in the standard predictor during high-attrition scenarios. In these environments, our method achieves an average improvement in mean returns of more than 20% and reduces performance variance by a mean of 64.7% compared to the standard unweighted baseline.
CommentsAccepted to 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)