JEPA-WAM:机器人操作中世界-动作模型的阶段级联合嵌入预测
StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
浏览论文内容
中文总结 AI 辅助
该研究针对通用机器人策略未明确建模任务阶段级未来的问题,提出JEPA-WAM模型,在50个RoboTwin 2.0任务上实现90.25%的整体成功率,成功减少了平均执行步骤数。
中文摘要 AI 辅助
通用机器人策略旨在将多模态观测结果和语言任务指令映射到不同任务中的动作,但现有方法通常将未来表示为固定的短时长视频-动作块,这种短期未来捕捉动作执行的局部场景演化,但未明确描述任务从当前阶段推进到下一阶段的阶段级未来。因此,我们区分机器人操作中两种互补的未来:用于捕捉局部场景演化的短期物理未来,以及用于表示任务进展的阶段级语义未来。我们提出JEPA-WAM,它在基于Motus的世界-动作模型(WAM)基础上,加入了阶段级联合嵌入预测架构(JEPA)预测器Stage-JEPA。给定当前观测和任务指令,Stage-JEPA使用冻结的V-JEPA2编码器提取当前状态表示,并预测下一推断阶段的潜在目标。在干净环境和随机环境中的50个RoboTwin 2.0任务上,JEPA-WAM的整体成功率达到90.25%,与最强基线相比,成功执行回合的平均执行步骤数减少了5.97%。
英文摘要
Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk. This short-term future captures local scene evolution for action execution, but it does not explicitly describe the stage-level future that specifies how a task should progress from its current stage to the next. We therefore distinguish two complementary futures for robot manipulation: a short-term physical future to capture local scene evolution and a stage-level semantic future to represent task progress. We introduce StageWAM, which augments a Motus-based World Action Model (WAM) with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor. Given the current observation and task instruction, Stage-JEPA uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage. Across 50 RoboTwin 2.0 tasks in clean and randomized environments, StageWAM achieves 90.25% overall success and reduces the mean number of execution steps in successful rollouts by 5.97% relative to the strongest baseline.