arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CIG-RL:面向不确定环境下源项估计的好奇心驱动信息引导强化学习

CIG-RL: Curiosity-Driven Information-Guided Reinforcement Learning for Source Term Estimation in Uncertain Environments

Junhee Lee, Seunghwan Kim, Hongro Jang, Hyungjin Kim, Hyoungho Park, Changseung Kim, Hyondong Oh

arXiv 2608.30673首次发表:更新:

发表机构

Korea Advanced Institute of Science and Technology (KAIST); Ulsan National Institute of Science and Technology (UNIST)(韩国科学技术院(KAIST); 蔚山国家科学技术研究院(UNIST))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对不确定环境下源项估计的问题,提出CIG-RL方法,通过好奇心驱动的信息引导强化学习,结合不确定性自适应主动感知奖励提升鲁棒性,经模拟与真实实验验证其有效性。

AI 中文摘要

源项估计(STE)旨在估计气体源的关键属性,对识别有害气体释放至关重要。信息论方法已被用于移动传感器的自主STE,因其在嘈杂环境中具有鲁棒性,但在线动作选择会产生大量计算成本。深度强化学习(DRL)凭借快速决策能力提供了一种有前景的替代方案,在基于DRL的STE中,智能体根据从嘈杂测量序列更新的源项信念状态选择动作。然而,现有方法依赖随机探索或仅依赖信念不确定性降低,而在DRL中缺乏有效探索策略,这可能会限制嘈杂环境下策略的鲁棒性。为解决该问题,我们提出一种好奇心驱动的信息引导强化学习,用于实现鲁棒且高效的STE。所提方法促进对训练期间未充分探索的新颖信念状态转移的主动探索,我们进一步引入一种不确定性自适应主动感知奖励,以在不确定性下引导高效的源搜索。高噪声条件下的模拟和真实世界实验证明了所提框架的鲁棒性和可行性,凸显了其在实际STE问题中的潜力。

英文摘要

Source term estimation (STE), which aims to estimate key properties of the gas source, is essential for identifying hazardous gas releases. Information-theoretic approaches have been adopted for autonomous STE using mobile sensors due to robustness in noisy environments, yet their online action selection incurs substantial computational cost. Deep reinforcement learning (DRL) provides a promising alternative with its fast decision-making capability. In DRL-based STE, the agent selects actions based on belief states of the source term updated from noisy measurement sequences. However, existing methods rely on random exploration or solely on belief uncertainty reduction without an effective exploration strategy in DRL, which can limit policy robustness in noisy environments. To address this, we propose a curiosity-driven information-guided reinforcement learning for robust and efficient STE. The proposed method promotes active exploration of novel belief state transitions that have not been sufficiently explored during training. We further introduce an uncertainty-adaptive active perception reward to guide efficient source search under uncertainty. Simulations under high-noise conditions and real-world experiments demonstrate the robustness and feasibility of the proposed framework, highlighting its potential for practical STE problems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑