发表机构
University of Science and Technology of China; Shanghai Artificial Intelligence Laboratory; East China University of Science and Technology; Zhejiang University; Northwestern Polytechnical University; TeleAI(中国科学技术大学; 上海人工智能实验室; 华东理工大学; 浙江大学; 西北工业大学; TeleAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出WOLF框架,利用循环世界模型预测未来观测并生成预测前沿,引导无人机激光雷达探索,显著提升效率和覆盖率。
AI 中文摘要
基于激光雷达的无人机(UAV)探索通过持续选择下一个观测位置来构建地图。然而,基于已测量地图的决策对遮挡物后方的空间延续性缺乏预见性,导致潜在的信息丰富方向未被识别。我们提出了WOLF,一个由世界模型引导的框架,通过预测未来观测来增强自主探索。在训练阶段,循环世界模型从探索轨迹中学习观测动态,循环记忆保留了解释连续视图中部分观测所需的空间上下文。基于该上下文,模型在探索过程中将观测历史与候选运动相结合,以预测局部占用率和可见性。为了进一步引导感知,预测前沿生成机制随后利用置信度、分支一致性和观测质量来对齐并融合这些预测,以识别有前景的区域。由此产生的预测前沿与已测量前沿相结合,用于引导几何视点选择和轨迹生成,而新的扫描则更新后续预测。在仿真中,我们的方法在Garage场景下以相当的覆盖率将平均终止时间相对于EPIC降低了10.9%,并在Tunnel场景中将平均覆盖率从42.12%提升至98.35%。真实世界实验进一步展示了在物理飞行过程中对学习模型进行机载部署以实现在线推理。
英文摘要
LiDAR-based unmanned aerial vehicle (UAV) exploration builds maps by continually selecting where to observe next. However, decisions based on the measured map provide limited foresight into spatial continuations behind occlusions, leaving potentially informative directions unrecognized. We present WOLF, a world-model-guided framework that predicts future observations to enhance autonomous exploration. In the training stage, a recurrent world model learns observation dynamics from exploration trajectories, with recurrent memory retaining the spatial context needed to interpret partial observations across successive views. Building on this context, the model combines observation history with candidate motions during exploration to predict local occupancy and visibility. To guide further sensing, a predictive frontier generation mechanism then aligns and fuses these predictions using confidence, branch agreement, and observation quality to identify promising regions. The resulting predictive frontiers join measured ones to guide geometric viewpoint selection and trajectory generation, while new scans update subsequent predictions. In simulations, our method reduces mean terminal time by 10.9% relative to EPIC in Garage at comparable coverage and increases mean coverage from 42.12% to 98.35% in Tunnel. Real-world experiments further demonstrate onboard deployment of the learned model for online inference during physical flight.