发表机构
Fordham University; City University of Hong Kong; IBM Research(福特汉姆大学; 香港城市大学; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出CASTLE框架,利用反事实语义-社会世界模型为独立多智能体强化学习提供在线上下文指导,在多个基准任务上超越最强基线。
AI 中文摘要
完全去中心化的多智能体强化学习(MARL),也称为独立学习,要求每个智能体仅使用其局部信息和经验进行学习和行动,无需集中式评论家或智能体间通信。这种严格的信息结构使得传统的奖励信号变得模糊不清。较差的回报可能源于无效的自身动作、不兼容的队友响应或有效的对手响应,但标量奖励本身无法揭示哪种解释是原因。我们认为,智能体可以通过前瞻性地比较候选动作的后果来更有效地学习,而不是仅从已实现的回报中诊断失败。我们引入了CASTLE(用于去中心化MARL中局部执行的反事实动作条件语义令牌),这是一个离线训练、在线上下文指导框架,包含两个互补的世界模型。一个局部动力学世界模型,在智能体的局部轨迹上离线预训练,总结了智能体的局部轨迹动力学和部分可观测性;而一个语义-社会世界模型预测每个候选自身动作的紧凑短期任务和社会后果。后者通过反事实模拟器回放进行训练,这些回放暴露了从相同记录的回放状态出发的替代动作下可能的队友和对手响应。在线学习和执行期间,两个世界模型保持冻结,智能体仅使用局部可用信息进行查询。它们的预测逻辑为独立的PPO策略提供上下文指导。在基准多粒子环境中的Tag、Spread和Adversary任务上,跨越30个匹配的种子,我们提出的CASTLE在评估方法中取得了最高的平均最终得分,分别超过每个任务的最强基线10.67、6.46和0.33个归一化点。
英文摘要
Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models. A Local Dynamics World Model, offline pre-trained over agents' local trajectories, summarizes the agent's local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action. The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state. During online learning and execution, both world models remain frozen and are queried by agents using only locally available information. Their prediction logits provide in-context guidance to an independent PPO policy. Across 30 matched seeds on Tag, Spread, and Adversary in the benchmark multi-particle environments, our proposed CASTLE achieves the highest mean final score among the evaluated methods, exceeding the strongest baseline on each task by 10.67, 6.46, and 0.33 normalized points, respectively.