空间记忆智能体:用于空间智能的基于经验的过程记忆
Spatial Memory Agent: Experience-Grounded Procedural Memory for Spatial Intelligence
浏览论文内容
中文总结 AI 辅助
该研究提出SMA框架,让冻结VLM智能体无需外部空间工具,通过无参数更新自进化提升空间推理,在多基准和模型上表现最优,提供了空间自进化的实用路径。
中文摘要 AI 辅助
空间智能正成为具身智能体、机器人规划和多模态助手的基础。为提升视觉语言模型(VLM)智能体的空间推理能力,现有研究主要遵循两条路线:一条采用后训练方法,如监督微调与强化学习;另一条采用智能体范式,即模型调用外部空间工具(如深度估计和三维重建工具)以收集中间空间证据。本文研究一条互补且被忽视的路径:在推理时不依赖外部专家空间工具,仅通过无参数更新的自进化,能否让冻结的VLM智能体提升其空间推理能力?我们提出空间记忆智能体(Spatial Memory Agent,SMA),这是一个基于经验的运行时框架,可将已验证的空间经验转化为可复用、可迁移的教训。在可验证的空间环境中,SMA查询冻结的VLM以获取预测答案和奖励,并利用验证器引导的反思从空间经验中提炼紧凑的可迁移教训;SMA还为每个教训分配迁移可靠性分数(Transfer Reliability Score,TRS),该分数初始化为均匀值,并根据后续检索结果进行校准,作为未来迁移可靠性的访问证据。在只读部署阶段,SMA通过语义过滤与相似度-TRS结合排序的方式检索教训,使检索到的记忆指导冻结模型推理。在五个代表性空间基准和四个基础VLM上,SMA在每个基础模型组中均达到最高宏平均,且在20次评估中的大多数评估中取得最佳准确率,为所评估的冻结模型规模和环境建立了一条实用的无参数更新空间自进化路径。
英文摘要
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLMs, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning through \textbf{parameter-update-free self-evolution}, without depending on external expert spatial tools at inference time? We present \textbf{Spatial Memory Agent (SMA)}, an experience-grounded runtime memory framework that converts verified spatial experience into reusable transferable lessons. Specifically, SMA first queries the frozen VLM in a verifiable spatial environment, obtains a predicted answer and reward, and uses verifier-guided reflection to distill compact transferable lessons stored in memory cards. SMA further assigns each memory card a \textbf{Transfer Reliability Score (TRS)}, which is initialized uniformly and calibrated from later retrieval outcomes as visit evidence of future transfer reliability. During read-only deployment, SMA retrieves memory cards through semantic filtering and combined similarity--TRS ranking, allowing the retrieved memory to guide frozen model inference. Experiments across five representative spatial benchmarks show that SMA achieves the best macro-average accuracy for all four base VLMs and the best accuracy in most individual evaluations, establishing a practical parameter-update-free path for spatial self-evolution through reusable experience.
发表机构
- Zhejiang University(浙江大学)
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。