基于仅三个距离传感器的无人机学习型动态避障方法
Learning-Based Dynamic Obstacle Avoidance for a UAV Using Only Three Range Sensors
浏览论文内容
中文总结 AI 辅助
提出一种基于深度强化学习(PPO)的无人机在线运动规划框架,仅用三个距离传感器,通过行为栅格图实现动态避障,仿真与实验均优于PPO变体和MPC,成功率显著提升。
中文摘要 AI 辅助
我们提出了一种基于学习的方法,用于在未知动态环境中以固定高度运行的无人航空器(UAV)的动力学在线运动规划,其中必须在极端部分可观测条件下实现对静态和动态障碍物的实时避让。该无人机仅通过一个自由度(偏航)进行控制,导致其运动受到约束且具有非完整性,类似于固定翼平台。所提出的框架将行为栅格地图表示与深度强化学习(DRL)相结合,使用近端策略优化(PPO)算法在连续控制中实现稳定的策略学习。关键思想是共同设计状态表示和控制策略,使得仅使用三个低成本的定向距离传感器即可实现可靠的导航,而无需依赖诸如激光雷达或基于视觉的系统等密集传感模态。行为栅格地图将稀疏测量动态聚合为结构化的局部表示,以支持用于避障和到达目标的实时决策。在具有不同大小和障碍物密度的环境中进行的大量仿真表明,所提出的标准方法和增强方法比PPO变体和模型预测控制(MPC)实现了更高的成功率(在小规模高拥堵场景中为94%对79–90%,在大规模高拥堵场景中为83%对62–71%),同时保持实时性能。在四个场景中的真实世界实验进一步证实了实际可行性,在测试条件下具有一致的目标到达行为且无碰撞。
英文摘要
We present a learning-based approach to kinodynamic online motion planning for an Unmanned Aerial Vehicle (UAV) operating at a fixed altitude in unknown dynamic environments, where real-time avoidance of both static and dynamic obstacles must be achieved under conditions of extreme partial observability. The UAV is controlled with a single degree of freedom (yaw only), resulting in constrained, nonholonomic motion similar to fixed-wing platforms. The proposed framework integrates a behavior grid map representation with Deep Reinforcement Learning (DRL), using Proximal Policy Optimization (PPO) for stable policy learning in continuous control. The key idea is the co-design of a state representation and control policy that enables reliable navigation using only three low-cost directional range sensors, without reliance on dense sensing modalities such as LiDAR or vision-based systems. The behavior grid map dynamically aggregates sparse measurements into a structured local representation that supports real-time decision-making for obstacle avoidance and target reaching. Extensive simulations across environments of varying sizes and obstacle densities demonstrate that the proposed standard and enhanced methods achieve higher success rates than PPO variants and Model Predictive Control (MPC) (94\% vs. 79--90\% in small-scale high-congestion scenarios, and 83\% vs. 62--71\% in large-scale high-congestion scenarios), while maintaining real-time performance. Real-world experiments across four scenarios further confirm practical feasibility, with consistent target-reaching behaviour and no collisions under the tested conditions.
发表机构
- SYSTEC-ARISE Research Center for Systems and Technologies, Faculty of Engineering, University of Porto(波尔图大学工程学院SYSTEC-ARISE系统与技术研究中心)
机构由 AI 辅助整理,请以论文原文为准。