发表机构
Institute for AI Industry Research (AIR), Tsinghua University; Z-Trans AI(清华大学人工智能产业研究院(AIR); Z-Trans AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Zeva框架,通过上下文因果学习让机器人从自身物理交互经验中学习,无需更新策略模型,在模拟与真实操作中性能优于前沿模型,还能自进化并跨任务泛化。
AI 中文摘要
仅通过预训练难以实现泛化性具身操作,原因在于现实世界中存在未见过的物理条件。我们认为机器人在现实部署过程中需要根据自身的物理交互实时学习,并利用这些知识指导后续动作。我们提出Zeva,这是首个能在保持策略模型冻结的同时,让机器人从自身物理交互经验中进行上下文学习的框架。Zeva采用因果交互提取器,将执行的动作及其引发的状态变化编码为因果交互信号,并存储在双时间尺度因果记忆中。对于后续动作,Zeva会从记忆中检索相关的因果交互信号,并将其作为上下文注入冻结的策略模型中。在模拟和真实操作环境中的实验表明,Zeva在对比的前沿视觉语言动作模型(VLAs)和腕臂机械手(WAMs)中取得了最佳性能,更重要的是,它能在部署过程中无需梯度更新实现自进化,其成功率会随着机器人积累交互经验而持续提升,且获得的交互经验可跨任务泛化。
英文摘要
Generalizable embodied manipulation remains difficult to achieve through pretraining alone, due to unseen physical conditions in the real world. We argue that robots need to learn from their own physical interactions on the fly during real-world deployment and use this knowledge to inform subsequent actions. We present Zeva, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen. Zeva employs a Causal Interaction Extractor to encode an executed action and its induced state change into a causal interaction signal, which is stored in a dual-timescale causal memory. For subsequent actions, relevant causal interaction signals are retrieved from memory and injected into the frozen policy model as context. Experiments in simulation and real-world manipulation demonstrate that Zeva achieves the best performance among the compared frontier VLAs and WAMs and, more importantly, enables self-evolution during deployment without gradient updates. Its success rate continues to improve as the robot accumulates interaction experience. Furthermore, the acquired interaction experience can generalize across tasks.