Vid2WAM:将视频扩散先验蒸馏为世界动作模型
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
浏览论文内容
中文总结 AI 辅助
Vid2WAM是将视频扩散先验蒸馏为WAM的离线框架,通过双监督通道与源感知残差动作适配,提升了机器人策略的新任务泛化性与数据效率,且推理高效。
中文摘要 AI 辅助
世界动作模型(World Action Models, WAMs)通过联合建模未来视觉动态与动作,改进机器人策略学习,但它们的可扩展性与泛化性仍受限于对成本高昂的专家演示的依赖。本文提出质疑:WAM的未来监督是否必须来自目标任务的专家轨迹?我们提出Vid2WAM,这是一个离线蒸馏框架,将大型视频基础模型的视觉扩散先验迁移到紧凑的WAM学生模型中。给定观测结果与语言指令,Vid2WAM通过两个互补通道进行监督蒸馏:任务条件下的未来展开直接监督学生模型的未来预测分支,而逆动力学模型则恢复特定于 embodiment 的伪动作以用于动作学习。为了稳健地整合合成与真实监督,我们引入了源感知残差动作适配,该方法学习共享动作主干周围的源特定修正,以减轻嘈杂伪动作的干扰。推理阶段,视频教师与逆动力学模型均被舍弃,仅保留WAM学生模型以实现高效部署。仿真与真实世界实验表明,在有限专家演示下,Vid2WAM提升了新任务泛化性与数据效率,同时保持低延迟推理。
英文摘要
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by their reliance on costly expert demonstrations. We challenge this by asking whether future supervision for WAMs must originate from target-task expert trajectories. In this paper, we propose Vid2WAM, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student. Given an observation and language instruction, Vid2WAM distills supervision through two complementary channels: task-conditioned future rollouts directly supervise the student's future prediction branch, while an inverse dynamics model recovers embodiment-specific pseudo-actions for action learning. To robustly integrate synthetic and real supervision, we introduce source-aware residual action adaptation that learns source-specific corrections around a shared action backbone and mitigates interference from noisy pseudo-actions. During inference, both the video teacher and inverse dynamics model are discarded, leaving only the WAM student for efficient deployment. Simulation and real-world experiments demonstrate that Vid2WAM improves novel-task generalization and data efficiency under limited expert demonstrations while preserving low-latency inference.