发表机构
Beijing Institute of Technology; LimX Dynamics(北京理工大学; LimX Dynamics)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ACG-WAM,通过动作条件几何预测辅助目标,联合学习视觉预测与机器人动作,在模拟和真实任务中均优于现有方法。
AI 中文摘要
世界动作模型联合学习视觉预测和机器人动作,提供了一种利用场景演化观测进行策略学习的方式。然而,其视频和动作损失并未为演示动作序列的几何后果提供明确目标。此外,时间注意力后提取的视觉特征可能包含未来观测,因此不适合作为辅助预测器的唯一当前视觉输入。我们提出了ACG-WAM及其辅助目标——动作条件几何联合嵌入预测架构(ACG-JEPA),该架构从当前观测和中间动作预测多个时间范围的几何特征,使用冻结的VGGT编码的每个当前和未来图像对的未来槽作为目标。我们将来自头部和腕部摄像头的监督应用于共享视觉嵌入,在时间混合之前,并在该URL处移除教师和辅助模块。在50个RoboTwin 2.0任务中,ACG-WAM在干净场景中达到93.46%的成功率,在随机化设置中达到最佳成功率(92.68%),并在两种设置下平均成功率为93.07%,优于对比方法;在真实机器人的三个任务中,它达到85.00%的成功率和91.67%的部分完成分数,分别超过Motus 10.00和9.17个百分点。代码:该URL。
英文摘要
World action models jointly learn visual predictionand robot actions, providing a way to use observations ofscene evolution for policy learning. Their video and actionlosses, however, provide no explicit target for the geometricconsequences of a demonstrated action sequence. Moreover,visual features taken after temporal attention can contain futureobservations, making them unsuitable as the sole current visualinput to an auxiliary predictor. We introduce ACG-WAMand its auxiliary objective, the Action-Conditioned GeometricJoint-Embedding Predictive Architecture (ACG-JEPA), whichpredicts geometric features at several horizons from the currentobservation and intervening actions, using the future slot of afrozen VGGT encoding of each current and future image pairas the target. We apply this supervision from the head and wristcameras to a shared visual embedding before temporal mixing,and remove the teacher and auxiliary modules at inference.On 50 RoboTwin 2.0 tasks, ACG-WAM achieves 93.46%success in clean scenes, with the best randomized success(92.68%) and mean across both settings (93.07%) among thecompared methods; across three tasks on a real robot, itachieves 85.00% success and 91.67% partial completion score,exceeding Motus by 10.00 and 9.17 percentage points, respec-tively. Code:https://github.com/RoboOpus/ACG-WAM.Website:https://RoboOpus.github.io/ACG-WAM.