发表机构
Stanford University; NVIDIA; University of Michigan, Ann Arbor(斯坦福大学; 英伟达; 密歇根大学安娜堡分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
T$^2$Mem通过测试时训练在预训练视觉-语言-动作策略中编码观测历史为快速权重,无需外部记忆,在16个RoboMME任务中将平均成功率从17.93%提升至56.83%,并实现至少3倍推理加速。
AI 中文摘要
依赖记忆的机器人操作要求策略能够利用在当前观测中不再可用的信息。仅保留历史记录是不够的:记忆必须保存支持未来行动的信息。一个挑战是,一个无记忆的基础模型能否仅从动作示范中学习保留和利用历史信息,而无需外部记忆支持。我们提出T$^2$Mem,一个在预训练的视觉-语言-动作策略内开发这种能力的框架,无需外部推理模型或记忆特定标注。T$^2$Mem使用测试时训练,通过在线自监督更新将观测历史编码为紧凑的快速权重,避免重复处理完整历史。一个基于观测的接口提取视觉-语言信息用于记忆形成,并将检索到的上下文提供给动作专家。动作监督塑造记忆学习保留和利用的内容,而交替的记忆-策略学习在优化过程中为每个组件提供固定的对应部分。在16个RoboMME任务中,T$^2$Mem将平均成功率从17.93%提高到56.83%,相对于无记忆的基础策略,并优于基准中报告的循环记忆方法,而受控的性能分析表明,相对于显式方法,推理速度至少提升3倍。项目网站:此https URL
英文摘要
Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T$^2$Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T$^2$Mem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, T$^2$Mem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods. Project website: https://yzliu84.github.io/T2MEM-project/