RoboTTT:机器人策略的上下文扩展
RoboTTT: Context Scaling for Robot Policies
- NVIDIA(英伟达公司)
- Stanford University(斯坦福大学)
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究提出RoboTTT,将测试时训练集成到机器人基础模型,通过序列动作强制和截断反向传播扩展视觉运动上下文到8K时间步长,解锁新能力,提升多任务性能,证明上下文长度是机器人基础模型新扩展轴。
AI中文摘要:
近期的机器人基础模型在单步或短历史视觉运动上下文下运行。我们引入了测试时训练机器人策略(RoboTTT),这是一种机器人模型和训练方法,可将视觉运动上下文扩展到8K时间步长,比现有技术策略高出三个数量级,且不增加推理延迟。在此上下文长度下,解锁了新的机器人能力,如从人类视频演示中一次性上下文模仿、即时策略改进、对扰动的鲁棒性以及在多阶段、长视野任务上更强的性能。还首次观察到随着预训练上下文长度增加,闭环性能稳步提升。核心是将测试时训练集成到机器人基础模型中,通过序列动作强制和截断反向传播来扩展训练上下文长度。在具有挑战性的真实机器人操作任务中,RoboTTT比单步上下文基线的整体性能提高了87%,并完全完成了五分钟、十阶段的装配任务,而基线模型从未做到。用8K时间步长上下文训练的RoboTTT比用1K时间步长预训练的相同模型性能高出62%,表明上下文长度是机器人基础模型新的扩展轴。
英文摘要:
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/