MARS-RA:基于具身多智能体协作中多模态比较的信用分配排序聚合
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation
浏览论文内容
中文总结 AI 辅助
本研究针对具身多智能体协作的信用分配难题,提出MARS-RA框架,将其转化为多模态模型生成的成对比较排序聚合问题,经理论与实验验证可引导智能体有效协作。
中文摘要 AI 辅助
信用分配是合作多智能体强化学习的核心挑战,尤其在具身AI场景中,存在反馈有限且延迟、活跃智能体数量动态变化的特点。我们提出MARS-RA框架,将信用分配重新表述为排序聚合问题,利用大型多模态模型生成的智能体间基于贡献的成对比较结果。这种从绝对估计到相对估计的转变,确保了对噪声和智能体动态参与的鲁棒性,将比较结果转化为基于势能的奖励塑造贡献分数。我们为该框架的收敛性和鲁棒性提供了理论证明,并表明Shapley值可作为可解释参考。在不同类型的挑战性任务上的实验结果表明,MARS-RA可引导智能体实现有效协作。
英文摘要
Credit assignment is a fundamental challenge in cooperative multi-agent reinforcement learning, particularly in embodied AI settings characterized by limited and delayed feedback as well as dynamically changing numbers of active agents. We propose MARS-RA, a framework that reformulates credit assignment as a rank aggregation problem using contribution-based pairwise comparisons among agents generated by large multimodal models. This shift from absolute to relative estimation ensures robustness against noise and dynamic agent participation, converting comparison results into contribution scores for potential-based reward shaping. We provide theoretical justification for the convergence and robustness of the proposed framework, and show that Shapley values can be used as an interpretive reference. Experimental results on challenging tasks of different types indicate that MARS-RA can guide agents toward effective cooperation.
发表机构
- Newcastle University(纽卡斯尔大学)
- University of Auckland(奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。