DroneWAM:面向无人机视觉导航的高效世界动作模型
DroneWAM: Efficient World Action Model for Drone Visual Navigation
浏览论文内容
中文总结 AI 辅助
提出DroneWAM,一种基于JEPA的高效世界动作模型,通过表示空间预测和自适应展开,在无人机视觉导航中实现高精度与低计算成本。
中文摘要 AI 辅助
世界动作模型为视觉导航智能体提供了一种方式,使其能够预测候选动作将如何改变未来观测,并根据预测的后果采取行动。对于无人机而言,这种能力必须在严格的精度和效率约束下运行。我们提出了DroneWAM,一种用于无人机视觉导航的高效世界动作模型。DroneWAM采用基于JEPA的架构,直接在表示空间中建模未来状态,避免了显式生成未来图像的计算成本。预训练的Resampler进一步将密集的编码器特征压缩为更少的潜在令牌,减少了每次想象步骤中重复的计算。我们还引入了自适应展开(adaptive rollout),其中经过偏好训练的Gate根据当前场景自适应地分配预测深度。为了支持在更丰富的空中运动下的学习,我们构建了DroneNav-6D,一个模拟视觉导航数据集,包含同步的RGB观测、6自由度飞行轨迹、控制命令和随机风扰动。在DroneNav-6D上,DroneWAM在比较方法中实现了最佳的轨迹精度。自适应展开进一步将平均预测深度从8降低到4.58,同时提高了轨迹精度,表明预测计算可以更有效地跨场景分配。代码和数据将公开发布。
英文摘要
World-action models give visual navigation agents a way to anticipate how candidate actions will change future observations and to act from the predicted consequences. For drones, this capability must operate under tight accuracy and efficiency constraints. We present DroneWAM, an efficient world-action model for drone visual navigation. DroneWAM adopts a JEPA-based architecture to model future states directly in representation space, avoiding the cost of explicit future image generation. A pretrained Resampler further compresses dense encoder features into fewer latent tokens, reducing the computation repeated at each imagined step. We also introduce adaptive rollout, where a preference-trained Gate adaptively allocates prediction depth according to the current scene. To support learning under richer aerial motion, we construct DroneNav-6D, a simulated visual navigation dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances. On DroneNav-6D, DroneWAM achieves the best trajectory accuracy among the compared methods. Adaptive rollout further reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy, demonstrating that predictive computation can be allocated more effectively across scenes. \href{https://github.com/1e12Leon/DroneWAM}{Codes and data} will be released.