发表机构
MemTensor (Shanghai) Technology Co., Ltd.; Renmin University of China; National University of Singapore; Shanghai Jiao Tong University; Tongji University(墨芯(上海)科技有限公司; 中国人民大学; 新加坡国立大学; 上海交通大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出首个记忆基础模型原型Metis,赋予基础模型原生记忆能力,通过新架构与优化目标实现,经实验验证其具备原生记忆能力并发布相关资源。
AI 中文摘要
人工智能智能体的近期进展已将原生能力日益内化为底层基础模型,催生了多模态基础模型与大型推理模型。然而,智能体记忆仍主要通过外部模块实现,原生记忆能力在很大程度上未被探索。本文在该方向迈出第一步,引入记忆基础模型,为基础模型赋予原生记忆能力。我们从两个视角形式化原生记忆:主干内持久且动态演化的记忆状态,以及通过模型计算自主存储和利用信息的原生记忆过程。我们表明原生记忆在架构、端到端优化和效率方面具有优势。基于此形式化,我们提出首个记忆基础模型原型Metis。Metis引入新架构,为基础模型配备原生记忆状态,使历史信息可被压缩入模型并通过记忆注意力访问。我们构建大规模记忆专用训练数据,并引入多个优化目标,通过中间训练获取这些原生记忆过程。Metis的在线记忆维护无需梯度,记忆更新仅需一次前向传播。推理时,所有已学习模型权重保持冻结,而原生记忆状态通过标准前向计算自主转换。通过大量实验,我们表明Metis具备原生记忆能力,并进一步对其优势、局限性和行为进行详细分析。为推动记忆基础模型的未来研究,我们发布项目及模型检查点。
英文摘要
Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external modules, leaving the native memory capability largely unexplored. In this paper, we take a first step toward this direction by introducing memory foundation models, which empower foundation models with native memory capabilities. We formalize native memory from two perspectives: a persistent and dynamically evolving memory state within the backbone, and native memory procedures that autonomously store and utilize information through model computation. We show that native memory offers advantages in architecture, end-to-end optimization, and efficiency. Based on this formulation, we propose Metis, the first prototype of memory foundation models. Metis introduces a new architecture that equips a foundation model with a native memory state, allowing historical information to be compressed into the model and accessed through memory attention. We construct large-scale memory-specific training data and introduce multiple optimization objectives to acquire these native memory procedures through mid-training. The online memory maintenance of Metis is gradient-free, and the memory update requires only a forward pass. At inference time, all learned model weights remain frozen, while the native memory states are autonomously transformed through standard forward computation. Through extensive experiments, we show that Metis exhibits native memory capabilities and further provide a detailed analysis of its strengths, limitations, and behaviors. To facilitate future research on memory foundation models, we release our project and model checkpoints.
Comments46 pages, 11 figures, 16 tables