arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniNav:用于视觉导航的统一世界-动作扩散模型

UniNav: A Unified World-Action Diffusion Model for Visual Navigation

Changqing Zhou, Yueru Luo, Zeyu Jiang, Changhao Chen

arXiv 2608.03244首次发表:更新:

AI 中文总结

UniNav是统一世界-动作扩散模型,通过单扩散过程生成未来视觉观测与航路点轨迹,其变体UniNav-Fast延迟0.1秒且精度无明显下降,在导航基准上ATE优于最强基线。

AI 中文摘要

图像目标视觉导航是具身智能体的基础能力。现有导航策略能高效预测航路点轨迹,但缺乏视觉前瞻能力;而导航世界模型可预测未来观测结果,但通常需要代价高昂的规划回滚。我们提出UniNav,一种统一的世界-动作模型,可通过单一扩散过程生成未来视觉观测结果和连续航路点轨迹。给定历史帧和目标图像,UniNav在单个Transformer中联合去噪视觉和航路点标记,在共享框架中统一未来预测与动作生成。为提升空间定位能力,我们融入感知几何的相机标记。我们同时在带轨迹标注的导航数据和仅视频数据上训练,使模型能从无航路点标注的多样视频中获益。基于该统一框架,我们引入两个变体:UniNav-Full联合预测可解释的未来观测结果及其对应轨迹,UniNav-Fast在推理时移除未来图像标记以实现高效轨迹预测。在导航基准上的实验表明,UniNav在所有数据集的ATE指标上优于最强基线。通过单步推理,UniNav-Fast实现0.1秒的延迟且无显著精度下降。代码将被开源。

英文摘要

Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts. We present UniNav, a unified world-action model that generates future visual observations and continuous waypoint trajectories through a single diffusion process. Given history frames and a goal image, UniNav jointly denoises visual and waypoint tokens within a single transformer, unifying future prediction and action generation in a shared framework. To improve spatial grounding, we incorporate geometry-aware camera tokens. We also train on both trajectory-labeled navigation data and video-only data, enabling the model to benefit from diverse videos without waypoint annotations. Based on this unified framework, we introduce two variants: UniNav-Full jointly predicts interpretable future observations and their corresponding trajectories, while UniNav-Fast removes future-image tokens at inference for efficient trajectory prediction. Experiments on navigation benchmarks show that UniNav outperforms the strongest baseline in ATE across all datasets. With one-step inference, UniNav-Fast achieves a latency of 0.1s without a substantial accuracy drop. Code will be released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑