发表机构
The Chinese University of Hong Kong, Shenzhen; Alibaba Group; Shenzhen Loop Area Institute; Nanjing University; Lovart AI; Voyager Research, Didi Chuxing(香港中文大学(深圳); 阿里巴巴集团; 深圳河套学院; 南京大学; Lovart AI; 滴滴出行 Voyager Research)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长视频世界模型中长程空间记忆管理难题,提出空间记忆智能(SMI)框架,利用多模态大语言模型的理解能力,通过四种原子操作实现记忆管理,显著提升记忆稀疏性、生成稳定性和空间一致性。
AI 中文摘要
长视频生成和世界模型通过根据用户动作和历史记忆预测未来观测,在交互式娱乐和具身模拟中展现出强大潜力。然而,随着记忆序列变长且结构日益复杂,管理长程空间上下文变得愈发具有挑战性,这需要一种更智能、更系统的记忆管理策略。基于多模态大语言模型(MLLMs)不断进步的空间推理能力以及统一模型的更广阔愿景,我们提出了空间记忆智能(SMI),这是首个在长视频世界模型中系统性地采用理解模型进行空间记忆管理的框架。SMI引入了四种协调的原子操作:空间聚类、簇内稀疏化、动作感知检索和可靠性感知过滤。跨多个基线、基准和世界模型骨干的大量实验证明了SMI的有效性和泛化性,在记忆稀疏性、生成稳定性和空间一致性方面取得了全面改进。
英文摘要
Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.
Comments32 pages. Project page: https://spatial-memory-intelligence.github.io/