Zero-WAM:基于人类视频的上下文世界-动作建模用于开放式任务泛化
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
浏览论文内容
中文总结 AI 辅助
Zero-WAM是基于人类视频的因果视频-动作模型,通过提出自动流水线构建HumanGen数据集、引入IFP目标,在RoboTwin 2.0仿真7个未见任务上平均成功率达47.0%,实现机器人操作零样本跨任务泛化。
中文摘要 AI 辅助
零样本跨任务泛化是机器人学习中的核心挑战,策略需执行训练期间从未见过的操作任务。在大语言模型中,只需在上下文指定新任务即可执行,无需任何参数更新,这种上下文学习(ICL)将泛化问题转化为任务指定问题。为实现跨任务泛化,本文将该范式应用于机器人操作,认为操作的自然任务指定是人类视频:与语言不同,它提供了关于预期任务演变的丰富视觉线索。本文提出Zero-WAM,一种因果视频-动作模型,通过遵循上下文人类视频指导来执行未见任务。为解决任务丰富的人机配对数据稀缺问题,本文提出一种自动流水线,将任务采样的机器人轨迹转换为语义匹配的人类视频,生成HumanGen数据集,包含8600个任务下的74200个人机ICL对。对于模型训练,本文进一步引入上下文未来块预测(IFP)目标,以抑制从已见任务中学到的捷径,并迫使策略从视频提示中提取任务信息。在RoboTwin 2.0仿真中的7个未见任务上,Zero-WAM实现了47.0%的平均成功率,比最强的视频-动作基线绝对提升29.5个百分点。在真实世界评估中,它遵循人类视频指导,泛化到涉及多物体场景、长时程操作和细粒度插入的未见任务配置。
英文摘要
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.