arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Skytopia:基于动作条件潜在世界模型的单目无人机导航

Skytopia: Monocular Drone Navigation with Action-Conditioned Latent World Models

Yuhang Zhang, Rangya Zhang, Yujing Shang, Zhuoyuan Yu, Weiying Wang, Steven Yang, Qingsong Yan, Chao Yan, Mir Feroskhan

arXiv 2609.26007首次发表:更新:

发表机构

Nanyang Technological University; Autel US; XGRIDS(南洋理工大学; 道通智能美国公司; 深圳显扬科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于动作条件潜在世界模型的单目无人机导航策略Skytopia,通过前向与逆向目标训练并丢弃预测器,在仿真中超越基线并在真实无人机上成功部署。

AI 中文摘要

单目无人机导航要求仅凭单个前视摄像头在未见过的环境中到达目标,而该摄像头提供的深度和尺度线索极少。世界模型通过模拟观测在动作下如何演变来解决这一问题,但它们是为执行而构建的:预测在部署时生成,并在每个控制步骤反馈到动作生成中。我们认为,策略从世界模型中需要的不是预测本身,而是生成预测所需的表示:在飞行中,已执行的动作几乎解释了观测之间所有的变化,因此预测归结为在已知位移下对静态场景的重投影。为此,我们引入skytopia,一种基于动作条件潜在世界模型的策略,以及用于训练它的3D高斯泼溅平台。前向目标根据预期运动预测下一观测的表示,逆向目标则从预测的转变中恢复该运动。由于预测从不进入动作生成,预测器被丢弃,一个策略即可服务于点目标、图像目标和无目标导航。仿真实验表明,skytopia在全部三种规格下均优于所有基线,分别达到57.8%、66.0%和49.0%的成功率,而丢弃预测器可消除59.4%的推理成本。同一策略随后未经微调便部署到物理无人机上,并在室内、室外开阔和树林环境中到达目标。

英文摘要

Monocular drone navigation requires reaching a goal in an unseen environment from a single forward-facing camera, which offers few cues for depth and scale. World models address this by modelling how observations evolve under actions, but they are built to be executed: the prediction is produced at deployment and fed back into action generation at every control step. We argue that what a policy needs from a world model is not the prediction but the representation required to produce it: in flight the executed action explains almost all of the change between observations, so prediction reduces to reprojecting a static scene under a known displacement. We therefore introduce skytopia, a policy built on an action-conditioned latent world model, and the 3D Gaussian Splatting platform on which it is trained. A forward objective predicts the representation of the next observation from the intended motion, and an inverse objective recovers that motion from the predicted transition. Because the prediction never reaches action generation, the predictor is discarded and one policy serves point-goal, image-goal, and goal-free navigation. Simulation experiments show that skytopia outperforms every baseline under all three specifications, attaining 57.8%, 66.0%, and 49.0% success rate, while discarding the predictor removes 59.4% of the inference cost. The same policy is subsequently deployed on a physical drone without fine-tuning and reaches goals in indoor, open outdoor, and woodland environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑