发表机构
National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology; College of Systems Engineering, National University of Defense Technology; Aerospace Information Research Institute, Chinese Academy of Sciences; BNRist, Tsinghua University; Southeast University(国防科技大学数字智能建模与仿真全国重点实验室; 国防科技大学系统工程学院; 中国科学院空天信息创新研究院; 清华大学北京国家研究中心; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对城市无人机目标搜索中部分可观测与反应式规划局限,提出SearchWorld循环状态空间世界模型,结合显式空间记忆与价值引导想象,在UAV-ON上将成功率提升至23.8%,实现前瞻性规划。
AI 中文摘要
自主无人机(UAV)目标搜索涉及在部分可观测条件下的感知、决策和行动的闭环过程。城市环境带来了若干挑战:大面积搜索区域和狭窄的以自我为中心的视野限制了覆盖范围,密集的三维几何结构约束了安全运动,开放世界的指令要求在有干扰物的情况下识别特定目标。许多现有方法通过显式地图或记忆表示来缓解部分可观测性问题,但大多仍是被动反应式的,仅基于过去观测进行推理,而不显式预测未来状态。世界模型通过想象推演实现前瞻性推理。然而,图像生成式世界模型可能带来高推理延迟,而潜在世界模型在空间接地规划方面仍具挑战性。我们提出SearchWorld,一种循环状态空间世界模型,将显式空间记忆与价值引导的想象相结合。该模型维护鸟瞰图(BEV)探索和障碍物记忆,并解码任务感知的空间价值层以引导搜索。认知-行动网络利用这一学习到的空间价值先验,通过想象推演改进策略,而无需训练单独标量评论家。训练过程从世界模型学习推进到专家模仿和基于想象的探索细化。在UAV-ON上,SearchWorld将成功率提升至23.8%(最强已发表智能体为19.5%),并将oracle成功率提升至35.5%,同时在未见场景上保持鲁棒性(成功率为19.9%)。通过将想象植根于显式空间表示,SearchWorld使UAV智能体能够前瞻性规划而非被动反应。
英文摘要
Autonomous unmanned aerial vehicle (UAV) object search involves a closed loop of perception, decision-making, and action under partial observability. Urban environments pose several challenges: large search areas and narrow egocentric views limit coverage, dense 3D geometry constrains safe motion, and open-world instructions require identifying a specific target among distractors. Many existing methods mitigate partial observability through explicit maps or memory representations, yet remain largely reactive, reasoning over past observations without explicitly predicting future states. World models enable prospective reasoning through imagined rollouts. However, image-generating world models can incur high inference latency, while spatially grounded planning remains challenging for latent world models. We propose SearchWorld, a recurrent state-space world model that connects explicit spatial memory with value-guided imagination. The model maintains BEV exploration and obstacle memory and decodes a task-aware spatial value layer to guide search. A cognition-action network uses this learned spatial value prior to improve the policy through imagined rollouts, without training a separate scalar critic. Training progresses from world-model learning to expert imitation and imagination-based exploration refinement. On UAV-ON, SearchWorld improves the success rate to 23.8% (19.5% for the strongest published agent) and raises oracle success to 35.5%, while remaining robust on unseen scenes (19.9% success rate). By grounding imagination in explicit spatial representations, SearchWorld enables UAV agents to plan prospectively rather than react.
Comments10 pages,2 figures