发表机构
KAIST AI(韩国科学技术院人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有多智能体世界模型对细粒度具身交互探索不足的问题,提出ME-World模型,通过联合去噪多智能体自我流等方式提升共享世界一致性等指标,在多智能体数据上验证了其性能优势。
AI 中文摘要
自我中心世界模型会根据智能体的动作预测第一人称观测结果,但大多数研究聚焦于单一智能体。真实的具身场景往往涉及多个在共享环境中行动与交互的智能体,现有多智能体世界模型依赖于 locomotion(移动)、相机控制或离散指令等粗粒度动作,对细粒度具身交互的探索不足。我们将多智能体自我中心世界建模问题,形式化为多个智能体在共享世界中通过细粒度动作交互时的同步自我流生成任务,该任务需要满足跨视图动作一致性、共享环境一致性,以及交互诱导的状态更新的一致传播。我们提出了多智能体自我中心世界模型(Multi-agent Egocentric World Model, ME-World),该模型在共享 token 序列中联合去噪多个自我流,将每个流基于所有智能体的目标视图姿态进行条件化,并通过共享环境记忆来锚定生成过程。我们在真实和合成的多智能体数据上进行训练与评估,并引入了针对环境一致性、更新一致性和身份一致性的共享世界一致性指标。实验结果表明,与现有方法相比,ME-World 在共享世界一致性、动作控制、身份保留和视频质量方面均有所提升。
英文摘要
Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.