arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31531cs.LG

HySTAR: 用于协作多智能体强化学习中稳定信用分配的锚定超图

HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning

发表机构电子科技大学
查看机构详情
  • University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang

首次发表
浏览论文内容

中文总结 AI 辅助

HySTAR通过锚定重叠稀疏超图作为稳定分解基础,结合时空编码器,在部分可观测协作多智能体强化学习中解决结构目标漂移,提升信用分配性能。

中文摘要 AI 辅助

在部分可观测和共享奖励条件下的协作多智能体强化学习,需要将团队结果分配给个体智能体和高阶联盟。MAPPO风格的评论家将联合行为压缩为一个全局值,而动态重构分组拓扑的评论家则随着交互或活跃智能体的演变,改变从智能体和联盟到价值组件的映射。我们将这种不一致性称为结构目标漂移。我们引入了HySTAR,一个基于MAPPO的框架,将自适应表示学习与时间一致的高阶价值分解基础分离。HySTAR将一个重叠稀疏超图锚定作为均匀覆盖的分解支架,使用时空编码器表示物理和任务相关的交互,并结合时间和结构相关性来构建智能体特定的优势。在SMAC、GRF、Traffic Junction和MPE上的实验表明,相对于MAPPO风格、价值分解和动态分组基线,HySTAR持续改进。在最难的SMAC设置中,HySTAR相对于MAPPO实现了16.7%的相对增益,相对于HYGMA实现了15.6%的相对增益,在所有六个GRF场景中排名第一,相对于MAGIC将Traffic Junction收敛轮次减少了高达40.2%,并获得了最高的MPE回合奖励。受控拓扑、智能体死亡、邻域和参数分析支持锚定分解支架同时适应传播表示的好处。

英文摘要

Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this inconsistency as structural target drift. We introduce HySTAR, a MAPPO-based framework that separates adaptive representation learning from a temporally consistent high-order value-decomposition basis. HySTAR anchors an overlapping sparse hypergraph as a uniformly covered decomposition scaffold, uses a spatiotemporal encoder to represent physical and task-dependent interactions, and combines temporal and structural relevance to construct agent-specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE demonstrate consistent improvements over MAPPO-style, value-factorization, and dynamic-grouping baselines. On the hardest SMAC settings, HySTAR achieves relative gains of 16.7\% over MAPPO and 15.6\% over HYGMA, ranks first on all six GRF scenarios, reduces Traffic Junction convergence epochs by up to 40.2\% relative to MAGIC, and obtains the highest MPE episode rewards. Controlled topology, agent-death, neighborhood, and parameter analyses support the benefit of anchoring the decomposition scaffold while adapting the propagated representations.

↑