DeltaWAM:通过Delta令牌实现以变化为中心的视觉预见,用于高效的世界动作模型
DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model
AI总结:
针对像素空间世界模型重建整个未来场景的高计算成本问题,DeltaWAM以紧凑的delta令牌为预测单元,基于DeltaWorld自回归预测帧间变化,并由流匹配动作专家生成动作,在LIBERO上取得92.8%成功率,且推理高效。
AI中文摘要:
世界动作模型(WAMs)为机器人操作提供视觉预见能力,但像素空间模型会反复重建整个未来场景,导致高昂的计算成本和时空冗余。在物理操作中,连续帧通常共享大部分视觉上下文;动作策略需要预测的正是它们之间的变化。我们提出了DeltaWAM,一种以变化为中心的世界动作模型,将紧凑的delta令牌作为未来预测的基本单位。每个令牌是一个单一向量,编码连续密集DINO特征图之间的变化。DeltaWAM基于DeltaWorld(一种在大规模视频上预训练的潜在世界模型),自回归地预测每个未来帧的一个delta令牌。随后,一个流匹配动作专家以预测的转换和当前DINO特征(作为空间锚点)为条件,生成动作块。DeltaWAM在两个H100 GPU上训练了256个GPU小时,拥有7.25亿参数,在LIBERO上实现了92.8%的平均成功率。在LIBERO-Pro的程序化扰动下,它也表现出稳健的泛化能力。推理每个动作块耗时142.1毫秒,峰值内存为3.86 GB。
英文摘要:
World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the changes between them are what an action policy needs to anticipate. We introduce DeltaWAM, a change-centric WAM that makes a compact delta token the unit of future prediction. Each token is a single vector encoding changes between consecutive dense DINO feature maps. DeltaWAM builds on DeltaWorld, a latent world model pretrained on large-scale videos, to autoregressively predict one delta token per future frame. A flow-matching action expert then conditions on the predicted transitions and current DINO features, which serve as spatial anchors, to generate action chunks. Trained for 256 GPU hours on two H100 GPUs, DeltaWAM has 0.725B parameters and achieves 92.8% average success on LIBERO. It also shows robust generalization under procedural perturbations on LIBERO-Pro. Inference takes 142.1 ms per action chunk with 3.86 GB peak memory.