arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于智能体特定偏好的多智能体强化学习

Multi-Agent Reinforcement Learning via Agent-Specific Preference

Ni Mu, Yao Luan, Yiqin Yang, Qing-Shan Jia

arXiv 2608.08604首次发表:更新:

发表机构

Tsinghua University; Chinese Academy of Sciences; CFINS; BNRist; Institute for Embodied Intelligence and Robotics; Institute of Automation(清华大学; 中国科学院; CFINS; BNRist; 嵌入式智能与机器人研究所; 自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出多智能体偏好集成学习(MAGPIE),通过智能体特定偏好建模解决多智能体强化学习中全局奖励设计难题,经理论证明与实验验证,其性能可媲美奖励工程基线,为奖励工程不适用场景提供有效策略学习方案。

AI 中文摘要

多智能体强化学习(MARL)是求解复杂协作任务的强大框架,但高度依赖定义良好的全局奖励函数。设计此类奖励颇具挑战性,尤其是在异构智能体系统中,单一标量目标可能无法捕捉多样化行为。本文提出多智能体偏好集成学习(MAGPIE),通过智能体特定偏好建模应对这些挑战:每个智能体由专用专家通过偏好信号评估,无需全局评估。我们从理论上证明,优化这些分散偏好会收敛到纳什均衡策略。为将局部偏好整合为一致的全局目标,我们从偏好数据构建智能体特定奖励模型,并通过单调聚合机制将其组合;进一步证明,优化该聚合奖励模型等价于训练纳什均衡策略。在基准多智能体任务及顺序生产线任务上的大量实验表明,MAGPIE 性能可与奖励工程基线媲美,展现出在精确奖励工程不切实际的场景中促进策略学习的潜力。

英文摘要

Multi-agent reinforcement learning (MARL) is a powerful framework for solving complex collaborative tasks, but it relies heavily on well-defined global reward functions. Designing such rewards is challenging, especially in systems with heterogeneous agents, where a single scalar objective may fail to capture diverse behaviors. In this paper, we introduce Multi-AGent Preference-Integrated lEarning (MAGPIE), which addresses these challenges through agent-specific preference modeling. Each agent is evaluated by a dedicated expert through preference signals, eliminating the need for global evaluation. We theoretically prove that optimizing these decentralized preferences converges to a Nash equilibrium policy. To integrate local preferences into a coherent global objective, we construct agent-specific reward models from preference data and combine them via a monotonic aggregation mechanism. We further prove that optimizing this aggregate reward model is equivalent to training the Nash equilibrium policy. Extensive experiments on benchmark multi-agent tasks and a sequential production line task show that MAGPIE achieves performance comparable to reward-engineered baselines, demonstrating its potential to facilitate policy learning in scenarios where precise reward engineering is impractical.

CommentsThis article has been accepted for publication in IEEE Transactions on Automation Science and Engineering. This is the author's version, which has not been fully edited, and the content may change prior to final publication. \c{opyright} 2026 IEEE. All rights reserved, including rights for text and data mining and training of artificial intelligence and similar technologies

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑