arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04701cs.RO

测试时训练作为机器人策略的残差记忆

Test-Time Training as Residual Memory for Robot Policies

Haoxuan Wang, Gengyu Zhang, Ramana Rao Kompella, Gaowen Liu, Yan Yan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出TTT-RM,将测试时训练用作残差记忆,通过历史解码器重建残差来补充有限记忆库,在模拟和真实任务中显著提升长时程机器人操作性能。

中文摘要 AI 辅助

记忆对于长时程机器人操作至关重要,其中成功的动作可能依赖于当前观察中无法再恢复的过去事件。然而,随着情节变长,保留完整历史变得越来越昂贵,这为记忆增强策略带来了根本性的可扩展性挑战。现有方法通过将选定的过去观察存储在有限记忆库中,或将交互历史压缩为通过测试时训练(TTT)获得的固定大小参数化状态来解决这一挑战。然而,这些公式并未明确区分可以从策略当前上下文中恢复的历史信息与必须超越当前上下文持续存在的信息。我们引入了TTT-RM,将TTT重新用作残差记忆,利用快速权重不是通用地压缩历史,而是通过保留无法从策略当前上下文恢复的任务相关历史信息来补充有限记忆库。具体而言,TTT-RM学习一个历史解码器,从当前观察和检索的记忆中重建历史表示。由此产生的重建残差捕获了该上下文无法解释的内容,并作为TTT的学习目标。TTT慢权重针对此残差目标进行优化,以便在线快速权重更新随时间学习编码互补的历史信息。然后查询快速权重状态以产生残差记忆表示,该表示调节动作生成。在记忆密集型模拟基准和真实世界任务上的大量实验表明,TTT-RM在多种记忆设计中持续改进,优于各种基线,并支持在三分钟、八阶段收纳任务上的持续执行。

英文摘要

Memory is essential for long-horizon robotic manipulation, where successful actions may depend on past events that are no longer recoverable from the current observation. As episodes grow longer, however, retaining the full history becomes increasingly costly, creating a fundamental scalability challenge for memory-augmented policies. Existing approaches address this challenge by either storing selected past observations in a bounded memory bank or compressing interaction history into a fixed-size parametric state through Test-Time Training (TTT). Yet these formulations do not explicitly distinguish between historical information that can already be recovered from the policy's current context and information that must persist beyond it. We introduce TTT-RM, which repurposes TTT as Residual Memory, using fast weights not to generically compress history but to complement a bounded memory bank by preserving task-relevant historical information that cannot be recovered from the policy's current context. Concretely, TTT-RM learns a history decoder that reconstructs historical representations from the current observation and retrieved memory. The resulting reconstruction residual captures what this context fails to explain and serves as the learning target for TTT. The TTT slow weights are optimized against this residual target so that online fast-weight updates learn to encode complementary historical information over time. The fast-weight state is then queried to produce a residual memory representation that conditions action generation. Extensive experiments on memory-intensive simulation benchmarks and real-world tasks show that TTT-RM consistently improves across multiple memory designs, outperforms diverse baselines, and supports sustained execution on a three-minute, eight-stage stowing task.

发表机构

  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)
  • Cisco Research(思科研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑