AI 中文总结
本文提出工作区令牌,通过训练时VLM蒸馏显著信息到轻量级记忆,部署时无需循环推理,提升记忆密集型任务策略性能。
AI 中文摘要
复杂的机器人操作任务通常需要对过去事件和动作的长期记忆。由于以完整历史为条件会使策略容易产生虚假相关性并降低性能,许多策略记忆方法通过昂贵的循环内VLM查询来压缩历史信息,仅处理与任务相关的显著信息。在本文中,我们提出了一种替代方法,即在训练时进行计算密集型的VLM查询,以学习一种轻量级的潜在记忆,该记忆可以在部署时高效查询。我们称这种表示为“工作区令牌”,其训练过程包括:(1) 使用VLM识别完成任务所需的当前和历史信息,然后(2) 通过集合重建解码器损失将这些信息蒸馏到工作区令牌中。在仿真和硬件实验中,我们表明工作区令牌可以在部署时作为观测的直接替代,使策略能够解决记忆密集型任务,而无需在循环中进行VLM推理。有趣的是,我们发现工作区令牌不仅更轻量,而且能带来更好的策略性能。
英文摘要
Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many approaches to policy memory involve compressing historical information through expensive VLM queries in-the-loop to process only task-salient information. In this paper, we propose an alternative approach in which computationally intensive VLM queries are made during train-time to learn a lightweight latent memory that can be efficiently queried at deployment time. Our representation, which we call the workspace token, is trained by (1) using a VLM to identify current and historical information necessary for completing a task, then (2) distilling these into the workspace token using a set-reconstruction decoder loss. In both simulation and hardware, we show that the workspace token can be used as a drop-in replacement for observations during deployment, enabling policies to solve memory-intensive tasks without the need for VLM reasoning in-the-loop, in effect serving as a latent harness for distilling a stronger reasoning models ability to solve long-horizon tasks to a reactive robotic policy. We further demonstrate that the workspace tokens are not only more lightweight, but also lead to better policy performance compared to conditioning policies on explicit modalities like curated past image frames, motivating a latent approach to history curation and reasoning model harnesses more broadly.
Comments26 pages; CoRL 2026; 11 figures