发表机构
University of California, Merced(加州大学默塞德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对小型无人机在湍流环境中定位气体源的不适定问题,提出信息引导安全强化学习框架,结合EMGR规划器与SAC探索策略,通过KL散度监控混合决策,实现近80%定位成功率且零安全违规。
AI 中文摘要
使用小型无人机系统(sUAS)对逃逸性气体排放进行自主定位本质上是一个不适定的逆问题。在湍流大气边界层中,高度间歇性的标量浓度场违反了经典梯度导航的假设,导致基于数据的估计器遭受严重噪声和虚假局部极小值的影响。为解决这些挑战,我们提出了一种信息引导的安全强化学习框架,并在一个定制的、GPU加速的3D模拟环境中进行评估,该环境将欧拉风场求解器与拉格朗日烟羽扩散模型耦合。我们识别出确定性信息寻求规划器的一个关键脆弱性——格拉姆偏差,即智能体基于有缺陷的早期估计贪婪地行动,使估计器缺乏空间多样性。为系统性地打破这种退化,我们的架构将经典的经验可观测性格拉姆(EMGR)规划器与学习的软演员-评论家(SAC)探索策略相结合。一个确定性的元监督器通过Kullback-Leibler(KL)散度主动监控估计器可靠性,动态地混合确定性利用与学习探索,以引导sUAS进入高信息区域。通过渐进式课程训练,并受到严格执行的鲁棒控制屏障函数(RCBF)的保护,我们的强化学习框架在复杂的移动源上实现了近80%的定位成功率,大幅优于经典基线(约30%),同时确保零安全违规。
英文摘要
The autonomous localization of fugitive gas emissions using small Unmanned Aircraft Systems (sUAS) constitutes a fundamentally ill-posed inverse problem. In turbulent atmospheric boundary layers, highly intermittent scalar concentration fields violate the assumptions of classical gradient-based navigation, causing data-driven estimators to suffer from severe noise and spurious local minima. To address these challenges, we introduce an Information-Guided Safe Reinforcement Learning framework evaluated within a custom, GPU-accelerated 3D simulation environment coupling an Eulerian wind solver with a Lagrangian puff dispersion model. We identify a critical vulnerability in deterministic information-seeking planners - a Gramian bias where agents act greedily upon flawed early estimates, starving the estimator of spatial diversity. To systematically break this degeneracy, our architecture integrates a classical empirical observability Gramian (EMGR) planner with a learned Soft Actor-Critic (SAC) exploratory policy. A deterministic meta-supervisor actively monitors estimator reliability via Kullback-Leibler (KL) divergence, dynamically blending deterministic exploitation with learned exploration to steer the sUAS into high-information zones. Trained via a progressive curriculum and safeguarded by a strictly enforced Robust Control Barrier Function (RCBF), our RL framework achieves nearly 80% localization success on complex, mobile sources - drastically outperforming classical baselines (~30%) - while ensuring zero safety violations.
Comments8 pages, 5 figures, 2 tables. Submitted to the 2026 IEEE Conference on Decision and Control (CDC). Code: https://github.com/sachingirime/Info-guided-safe-RL