arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EmbodiedMemory-Bench:面向长时程具身任务的具身记忆基准测试

EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

Lizhou Liang, Xinyu Zhong, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Qinfeng Li, Peng Li, Jintao Chen, Xuhong Zhang, Wenqi Zhang

arXiv 2609.28236首次发表:更新:

发表机构

Zhejiang University; Central South University; Institute of Software, Chinese Academy of Sciences(浙江大学; 中南大学; 中国科学院软件研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程具身任务中智能体记忆能力不足的问题,提出EMem-Bench基准测试,包含四个任务族共2554个交互情节,并设计外部记忆系统EMem及8B策略模型EMem-8B,实验表明现有模型表现薄弱,而EMem显著提升性能。

AI 中文摘要

长时程具身交互要求智能体在观察、行动和应对变化的过程中,持续保留并更新关于环境的信息。然而,当前的智能体难以可靠地维持这种记忆。我们的分析将这一局限追溯到四个关键缺陷:细粒度视觉记忆薄弱、动态世界状态跟踪不可靠、未能记录由交互结果揭示的世界状态,以及从先前经验中泛化的能力有限。然而,现有基准测试并未在长时程具身交互过程中直接评估这些记忆能力。为弥补这一空白,我们引入了EmbodiedMemory-Bench(EMem-Bench),包含四个任务族中的2,554个交互式情节。EMem-Bench要求智能体从交互历史中构建和更新记忆,然后利用这些记忆通过在环境中行动来完成后续任务。我们进一步提出了Embodied-Memorizer(EMem),一种外部记忆系统,将具身经验组织为空间、事件和场景记忆。我们还训练了EMem-8B,一个管理并使用这些记忆的8B策略模型。我们评估了多种开源和专有MLLM以及具有代表性的多模态记忆系统。结果表明,当前模型在这四个挑战上表现仍然薄弱且不均衡。在匹配骨干网络的情况下,EMem在评估的记忆系统中取得了最佳整体性能,并提升了开源和专有模型的表现,而EMem-8B在其骨干基础上进一步提升。项目页面:https://zju-omniai.github.io/EmbodiedMemoryBench/

英文摘要

Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounter change. Yet current agents struggle to maintain such memory reliably. Our analysis traces this limitation to four key deficiencies: weak fine-grained visual memory, unreliable dynamic world-state tracking, failing to record world state revealed by interaction outcomes, and limited generalization from prior experience. However, existing benchmarks do not directly assess these memory capabilities during long-horizon embodied interaction. To address this gap, we introduce EmbodiedMemory-Bench (EMem-Bench), comprising 2,554 interactive episodes across four task families. EMem-Bench requires agents to build and update memory from interaction history, then use it to complete a later task by acting in the environment. We further present Embodied-Memorizer (EMem), an external memory system that organizes embodied experience into spatial, event, and scene memories. We also train EMem-8B, an 8B policy that manages and uses these memories. We evaluate a diverse range of open-source and proprietary MLLMs and representative multimodal memory systems. Results show that current models remain weak and uneven across the four challenges. Under matched backbones, EMem achieves the best overall performance among the evaluated memory systems and improves both open-source and proprietary models, while EMem-8B further improves over its backbone. Project page: https://zju-omniai.github.io/EmbodiedMemoryBench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑