发表机构
Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在动态杂波中四旋翼安全飞行问题,提出预期风险引导强化学习框架,利用特权模拟器状态构建风险地图,通过非对称架构训练网络自我预测风险,结合轻量级编码器提取线索,实验证明该方法有效提高安全裕度和飞行效率,且能实现模拟到现实的零样本转移。
AI 中文摘要
在杂乱和动态环境中的安全四旋翼导航不仅取决于即时几何感知,更关键的是要预测由相对运动引起的碰撞风险。传统模块化管道常受感知延迟影响,而依赖隐式标量奖励的端到端学习方法在无物理基础监督时难以提取可靠时空特征。为此,我们提出一个预期风险引导强化学习框架。利用特权模拟器状态,基于最近接近点构建定向对齐的未来碰撞风险地图。通过非对称演员-评论家架构,网络被训练自我预测这种结构化风险,在部署时明确指导视觉策略。一个轻量级时空编码器直接从机载深度序列提取运动线索,绕过显式目标跟踪或光流估计。大量模拟和真实世界实验表明,与现有基线相比,我们的方法有效提高了在密集动态杂波中的安全裕度和飞行效率。此外,学习到的策略在物理四旋翼上实现了强大的零样本从模拟到现实的转移,验证了我们方法的有效性及其从模拟到现实的强大泛化能力。
英文摘要
Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided reinforcement learning framework. Leveraging privileged simulator states, we construct a directionally aligned future collision risk map based on the Closest Point of Approach (CPA). Through an asymmetric actor-critic architecture, the network is trained to self-predict this structured risk, which explicitly guides the visual policy during deployment. A lightweight spatio-temporal encoder extracts motion cues directly from onboard depth sequences, bypassing explicit object tracking or optical flow estimation. Extensive simulated and real-world experiments demonstrate that our method effectively improves safety margins and flight efficiency in dense dynamic clutters compared to existing baselines. Furthermore, the learned policy achieves robust zero-shot Sim-to-Real transfer on a physical quadrotor, relying purely on abstracted spatio-temporal depth sequences and its self-predicted risk priors, validating the effectiveness of our approach and its robust generalization from simulation to reality.
Comments8 pages, 7 figures. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)