arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

部分可观测多智能体导航中零样本对手适应的分层信念建模

Hierarchical Belief Modeling for Zero-Shot Opponent Adaptation in Partially Observable Multi-Agent Navigation

Kowei Shih, Lu Cheng, Zeyu Wang, Yeyun Xu, Kejian Tong

arXiv 2609.12422首次发表:更新:

发表机构

Tsinghua University; Stevens Institute of Technology; University of California, Los Angeles; Texas A&M University(清华大学; 史蒂文斯理工学院; 加利福尼亚大学洛杉矶分校; 德克萨斯A&M大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对Lux AI第3赛季部分可观测导航问题,提出分层智能体HORIZON,结合对称感知、双记忆信念追踪、图注意力、信息增益探索及对手条件策略混合,经PPO训练,显著提升比赛胜率与适应能力。

AI 中文摘要

Lux AI 第3赛季要求智能体在部分可观测性、随机化的回合级动态以及五局三胜的比赛结构下行动,这种结构同时奖励战术执行和快速适应。我们提出HORIZON,一种分层智能体,它结合了对称感知的空间感知、双记忆信念追踪、以遗物为中心的图注意力、信息增益驱动的探索以及对手条件策略混合。HORIZON将短视界控制与跨比赛元推理分离,而辅助信念和世界模型目标稳定学习。使用PPO在大规模JAX模拟器中训练,所得智能体显式推断隐藏的游戏参数和对手风格。实验表明,与强循环和前馈基线相比,在比赛胜率、回合胜率、适应增益和联赛评分方面均有持续提升。

英文摘要

Lux AI Season 3 requires agents to act under partial observability, randomized episode level dynamics, and a best of five match structure that rewards both tactical execution and fast adaptation. We present HORIZON, a hierarchical agent that combines symmetry aware spatial perception, dual memory belief tracking, relic centric graph attention, information gain driven exploration, and an opponent conditioned policy mixture. HORIZON separates short horizon control from cross match meta reasoning, while auxiliary belief and world model objectives stabilize learning. Trained with PPO in a large scale JAX simulator, the resulting agent explicitly infers hidden game parameters and opponent style. Experiments show consistent gains in match win rate, episode win rate, adaptation gain, and league rating over strong recurrent and feed forward baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑