arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32313cs.RO

MemTransfer:具身决策中超越匹配经验的记忆基准测试

MemTransfer: Benchmarking Memory Beyond Matched Experience in Embodied Decision-Making

Haiming Tang, Xianjie Dai, Gujie Shao, Zuyi Guo, Jingguang Li, Kailang Ma, Yihong Tang, Heye Huang

首次发表
浏览论文内容

中文总结 AI 辅助

MemTransfer基准测试比较六种记忆表示在具身导航中的表现,发现对起始位姿或路线变化的鲁棒性不相互关联,且情景记忆在增加相关演示时提升显著,强调了评估存储信息与决策时应用的重要性。

中文摘要 AI 辅助

记忆使具身智能体能够复用过去的经验,然而保留有用信息并不能确保智能体在条件变化时能够应用这些信息。我们提出了MemTransfer,一个在共享的冻结视觉-语言模型策略下,比较六种记忆表示(一种工作记忆基线和五种过去经验表示)的基准测试。该基准包含模拟仓库中十种任务类型的一百个导航案例,并以专家演示作为历史数据。三项比较分别改变起始位姿、路线可用性以及历史数据的数量与任务相关性。在每任务一条演示的情况下,全上下文记忆和情景记忆在原始演示起点分别达到95.3%和100.0%的成功率,但在新的测试起点损失了48至49个百分点。摘要在这两个测试起点之间变化不大,然而在每任务四条演示的情况下,其在新测试起点保留的未变路线成功率(39.3%)低于工作记忆(44.8%)或两种轨迹记忆(56-58%)。在新测试起点,将相关演示从一条增加到四条使情景记忆的成功率提升14.3个百分点,而其他评估的表示提升不超过1.3个百分点。将一半的相关历史替换为其他任务的经历会降低两种轨迹记忆的成功率。这些结果表明,对一种不匹配的鲁棒性并不暗示对另一种不匹配的鲁棒性,从而激励对存储信息及其在决策时使用的评估。

英文摘要

Memory lets an embodied agent reuse past experience, yet retaining useful information does not ensure that the agent can apply it when conditions change. We present MemTransfer, a benchmark comparing six memory representations, a working-memory baseline and five representations of past experience, under a shared frozen vision-language-model policy. It comprises 100 navigation cases across ten task types in a simulated warehouse, with expert demonstrations supplying the history. Three comparisons vary the starting pose, route availability, and amount and task relevance of history. With one demonstration per task, Full-context and Episodic memory reach 95.3% and 100.0% success at the original demonstration start, but lose 48-49 percentage points at a new test start. Summary changes little between these two test starts, yet with four demonstrations per task it retains a smaller fraction of its unchanged-route success after blocking (39.3%) than Working memory (44.8%) or the two trajectory memories (56-58%). At the new test start, increasing from one to four relevant demonstrations raises Episodic success by 14.3 percentage points, while the other evaluated representations gain no more than 1.3 percentage points. Replacing half of the relevant histories with other-task experience lowers success for both trajectory memories. These results show that robustness to one kind of mismatch does not imply robustness to another, motivating evaluation of both stored information and its use at decision time.

↑