DiffWAM:一种快速高效的导航世界动作模型
DiffWAM: A Fast and Efficient Navigation World Action Model
浏览论文内容
中文总结 AI 辅助
DiffWAM提出一种几何条件化的导航世界动作模型,直接从冻结视频模型的预测特征中恢复无人机连续3D轨迹,无需未来视频合成,在DiffWAM-1000基准上实现0.3492米RMSE和74.40%成功率,并支持实时板载执行。
中文摘要 AI 辅助
预训练的视频基础模型为具身导航编码了丰富的语义和时空先验,然而将这些先验转化为无人机运动通常需要昂贵的未来视频合成和几何重建。我们研究是否可以直接从冻结视频模型的预测表示中恢复未来视觉预测中隐含的运动。为此,我们提出了DiffWAM,一种几何条件化的导航世界动作模型,它直接将多级预测特征转化为连续的相机轨迹。其Grid-Motion模块保留了时空运动关联,而Latent2Pose利用首帧几何将这些关联锚定,以恢复具有度量意义的3D运动。完整的视频展开和几何重建仅用于离线监督,从而在部署时消除了未来视频解码和多帧重建。我们进一步引入了FastDreamer,它将预测和几何计算与正在进行的飞行重叠,并执行时间戳感知的异步轨迹交接,以实现连续的无人机执行。DiffWAM在包含1,000个样本的DiffWAM-1000基准上实现了0.3492米的轨迹均方根误差和74.40%的端点成功率,而具有代表性的真实世界实验展示了包括受限穿越、环绕、S形飞行和多阶段导航在内的复杂行为。板载DiffWAM-Flash实现进一步在NVIDIA Jetson AGX Thor上达到了1.08秒的模型流水线延迟。这些结果表明,预测性视频表示可以有效地转化为连续的3D运动,为生成-然后重建的导航流水线提供了一种直接替代方案。项目页面:此https URL。
英文摘要
Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be recovered directly from the predictive representations of a frozen video model. To this end, we present DiffWAM, a geometry-conditioned navigation world-action model that directly transforms multi-level predictive features into continuous camera trajectories. Its Grid-Motion module preserves spatial-temporal motion associations, while Latent2Pose grounds them with first-frame geometry to recover metrically meaningful 3D motion. Complete video rollouts and geometric reconstruction are required only for offline supervision, eliminating future-video decoding and multi-frame reconstruction during deployment. We further introduce FastDreamer, which overlaps predictive and geometric computation with ongoing flight and performs timestamp-aware asynchronous trajectory handoff for continuous UAV execution. DiffWAM achieves a trajectory RMSE of 0.3492 m and an endpoint success rate of 74.40% on the 1,000-sample DiffWAM-1000 benchmark, while representative real-world experiments demonstrate complex behaviors including constrained traversal, orbiting, S-shaped flight, and multi-stage navigation. An onboard DiffWAM-Flash implementation further reaches 1.08 s model-pipeline latency on NVIDIA Jetson AGX Thor. These results demonstrate that predictive video representations can be efficiently grounded into continuous 3D motion, providing a direct alternative to generate-then-reconstruct navigation pipelines. Project page: https://zzmmzzm.github.io/diffwam.github.io/.
发表机构
- Zhejiang University(浙江大学)
- Differential Robotics(微分机器人公司)
机构由 AI 辅助整理,请以论文原文为准。