arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MEMOBench:用于机器人操作的过程级内存基准

MEMOBench: A Process Level Memory Benchmark for Robotic Manipulation

Haiyang Sun, Haoxiao Wang, Junming Chen, Weicheng Fang, Zihao Su, Jingkun Yi, Wenyou Yi, Hao Chen, Zhou Zhao

arXiv 2609.07047首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MEMOBench提出过程级内存基准,含30个任务和4,200个检查点,定义存储、更新、压缩率指标,发现现有VLA策略内存能力不足,并提供诊断与训练监督。

AI 中文摘要

机器人操作通常需要对不再可见的信息进行行动,然而视觉-语言-行动(VLA)策略通常在当前观察在很大程度上决定下一步行动时进行评估。现有的机器人内存基准暴露了这一差距,但它们仍然主要依赖于最终任务的成功,因此将遗忘与操作失败混为一谈。我们提出了MEMOBench,一个用于机器人操作中过程级内存评估的基准。MEMOBench包含30个历史依赖任务、1,500个专家演示以及来自84个模板的4,200个可执行检查点实例。每个检查点将粗粒度到细粒度的语言与模拟器谓词配对,并标记一种内存操作:存储、更新或压缩。这些注释定义了内存存储率、内存更新率和内存压缩率,这些指标在任务成功的同时衡量内存保真度。在标准和内存增强的VLA策略中,最强的内存模块基线仅达到31.9%的平均成功率,且高存储率常常与较弱的更新和压缩率共存。检查点语言还监督语义、对比和逐帧内存对齐目标,在不同内存操作上产生适度提升。MEMOBench为内存基础的机器人策略提供了诊断性评估套件和训练监督。项目页面可从此https URL访问。

英文摘要

Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Action policies are usually evaluated when the current observation largely determines the next action. Existing robotic memory benchmarks expose this gap, but they still rely mainly on final task success and therefore conflate forgetting with manipulation failure. We present \textbf{MEMOBench}, a benchmark for process level memory evaluation in robotic manipulation. MEMOBench includes 30 history dependent tasks, 1{,}500 expert demonstrations, and 4{,}200 executable checkpoint instances from 84 templates. Each checkpoint pairs coarse to fine language with a simulator predicate and labels one memory operation: Storage, Update, or Compression. These annotations define Memory Storage Rate, Memory Update Rate, and Memory Compression Rate, which measure memory fidelity alongside task success. Across standard and memory augmented VLA policies, the strongest memory module baseline reaches only 31.9\% average success rate, and high storage often coexists with weak update and compression. Checkpoint language also supervises semantic, contrastive, and framewise memory alignment objectives, yielding modest gains across different memory operations. MEMOBench provides a diagnostic evaluation suite and training supervision for memory grounded robotic policies. The project page is available at https://github.com/Collab-Gen/MEMOBench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑