发表机构
Beihang University; JD Future Academy(北京航空航天大学; 京东未来研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DyMD通过自适应重加噪和动力学引导的伪分数跟踪,在少步视频世界模型蒸馏中保留交互动力学,提升任务遵循度与规划成功率。
AI 中文摘要
大型视频扩散模型为具身预测和学习提供了强大的先验,但其多步采样对于交互式下游使用而言成本高昂。分布匹配蒸馏(DMD)能够实现少步视频生成,但在保持视觉质量的同时,可能会抑制机器人-物体运动。通过检查DMD的教师信号和伪分数信号,我们发现弱重新加噪会使教师后验集中在运动不足的轨迹附近,从而限制了恢复运动的指导。同时,较强运动的轨迹往往会产生更大的伪分数拟合误差,这可能阻碍生成器对交互动力学的学习。我们提出DyMD,一种DMD框架,它使教师监督和评论器拟合都适应不断演化的学生。时间亲和性条件重新加噪采样通过将基础调度与由局部后验变化驱动的教师先验混合,使时间步分布适应每个轨迹当前的交互保真度,从而平衡运动恢复和外观细化。为了更好地跟踪较强运动的轨迹,动力学引导的伪分数跟踪使用噪声条件预测器,根据潜在时间动力学估计噪声相对拟合难度,然后在评论器损失中对预测困难的轨迹进行加权。使用DyMD,我们将一个14B教师模型蒸馏为一个四步1.3B学生模型,推理时无需辅助模块。在具身视频基准上,与基础DMD相比,学生模型在R-Bench任务遵循度上提高了9.6个百分点,在PAI-Bench-G领域得分上提高了5.1个百分点,同时保持了相当的视觉质量。作为下游动作规划的骨干,我们的学生模型在两个WorldArena任务中实现了34%的平均成功率,而基础DMD为16%。
英文摘要
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. Meanwhile, stronger-motion rollouts tend to incur larger fake-score fitting errors, which can hinder the generator's learning of interaction dynamics. We propose DyMD, a DMD framework that adapts both teacher supervision and critic fitting to the evolving student. Temporal affinity--conditioned re-noise sampling adapts the timestep distribution to each rollout's current interaction fidelity by mixing the base schedule with a teacher prior motivated by local posterior variation, thereby balancing motion recovery and appearance refinement. To better track stronger-motion rollouts, dynamics-guided fake-score tracking uses a noise-conditioned predictor to estimate noise-relative fitting difficulty from latent temporal dynamics, then upweights predicted-hard rollouts in the critic loss. Using DyMD, we distill a 14B teacher into a four-step 1.3B student with no auxiliary modules at inference. On embodied-video benchmarks, the student improves R-Bench task adherence by $9.6$ percentage points and PAI-Bench-G Domain score by $5.1$ points over Base DMD while maintaining comparable visual quality. As a backbone for downstream action planning, our student achieves 34% mean success across two WorldArena tasks, compared with 16% for Base DMD.