一试即优:用于跨情节探索摊销的去冗余过程记忆
Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization
查看机构详情
- Tsinghua University(清华大学)
- DISCOVER Robotics(DISCOVER机器人公司)
- Northeast Agricultural University(东北农业大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究如何让机器人避免重复探测隐藏状态物体,提出面向实例的记忆框架IOM,用视觉语言模型实例化,在多任务中减少操纵操作,提升效率且不降低成功率,即使过程错误也能成功。
中文摘要 AI 辅助
操纵具有隐藏内部状态的物体,如带锁的微波炉,迫使机器人在行动前进行探测。然而,解决过一次某个实例的机器人,再次遇到该实例时会重新运行相同的探测,因为现有的跨情节记忆以任务成功为目标,围绕状态而非物体或重新探索的成本来组织重用。我们提出了面向实例的记忆(IOM),这是一个以物体为中心的框架来摊销这种探索:从一次揭示隐藏状态的单次遭遇中,无论是否成功,IOM记录操纵该实例的简短过程,根据物体的可识别特征对其进行索引,并将其作为对过程条件策略的软偏差注入。后续遭遇时能识别物体并回忆其过程而非重新探索。我们用现成的视觉语言模型(VLM)实例化这种提炼,无需特定任务训练就能将每次遭遇解析为过程。在四个铰接物体任务中,两个在模拟环境(微波炉、门),两个在真实机器人上(瓶子、柜子),一个神谕过程记忆在不降低成功率的情况下,比重新探索减少16 - 30%的操纵操作,VLM实例化能直接恢复69 - 88%的节省。由于过程是对反馈驱动策略的软偏差,即使检索到的过程错误也能成功,约12%的门实例就是如此。在所有任务中,好处纯粹是效率方面的:成功率从不降低,在真实机器人上甚至有所提高。代码将在接受后发布。
英文摘要
Manipulating objects with hidden internal state, such as a latched microwave, forces a robot to probe before it can act. Yet a robot that has solved an instance once re-runs the same probes whenever it encounters that instance again, because existing cross-episode memories target task success and organize reuse around states, not the object or the cost of re-exploring it. We present Instance-Oriented Memory (IOM), an object-centric framework that amortizes this exploration: from a single encounter that uncovers the hidden state, whether or not it succeeds, IOM records a short procedure for manipulating that instance, keys it on the object's identifiable features, and injects it as a soft bias on a procedure-conditioned policy. A later encounter recognizes the object and recalls its procedure instead of re-exploring. We instantiate this distillation with an off-the-shelf vision-language model (VLM) that parses each encounter into the procedure without task-specific training. Across four articulated-object tasks, two in simulation (microwave, door) and two on a real robot (bottle, cabinet), an oracle procedure memory cuts manipulation operations by 16-30% over re-exploration at non-regressing success, and the VLM instantiation recovers 69-88% of that saving out of the box. Because the procedure is a soft bias on a feedback-driven policy, an incorrect memory is recovered from rather than obeyed: success holds even when a retrieved procedure is wrong, as for $\approx$12% of door instances. Across all tasks the benefit is purely one of efficiency: success never regresses, and on the real robot even improves. Code will be released upon acceptance.