发表机构
University of California, Los Angeles; Adobe Research(加利福尼亚大学洛杉矶分校; Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多智能体交互世界模型共享状态维护难题,提出WorldWeaver模型,利用跨智能体世界状态寄存器及混合变压器设计,经双智能体我的世界视频生成实验验证,该模型能提升逻辑一致性与生成质量。
AI 中文摘要
多智能体交互世界模型不仅要生成一致的观测结果,还要维护跨智能体持久且跨视图演变的世界状态。现有的自回归视频扩散管道将观测历史作为条件上下文传递,这使得在多智能体和多视图设置中难以维护共享状态。我们提出了WorldWeaver(W^2),一种流式多智能体视频扩散模型,它通过跨智能体世界状态寄存器增强了展开过程:可学习的令牌存储共享世界信息、跟踪单个智能体状态,并在每个生成块后动态更新。我们用跨越单个智能体状态、包括鸟瞰图的全局状态视图和场景文本的监督信号来为这些寄存器提供基础。我们进一步用混合变压器设计改进了架构,该设计对世界状态建模和视觉帧建模使用单独的权重。在双智能体我的世界视频生成中的广泛实验表明,显式的世界状态建模提高了逻辑一致性和生成质量。
英文摘要
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk. We ground these registers with supervision signals spanning individual agent status, global state views including bird's-eye views, and scene text. We further improve the architecture with a Mixture-of-Transformers design that uses separate weights for world state modeling and visual frame modeling. Extensive experiments in two-agent Minecraft video generation show that explicit world-state modeling improves logical consistency and generation quality.
CommentsProject page: https://vail-ucla.github.io/worldweaver/