发表机构
National Engineering School of Tunis, University of Tunis El Manar; University of Tripoli; University of Picardie Jules Verne Amiens(突尼斯艾尔玛纳尔大学突尼斯国立工程师学校; 的黎波里大学; 亚眠皮卡第朱尔斯·凡尔纳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对稀疏奖励长视野强化学习挑战,提出两级分层强化学习框架,高级战略规划与低级SAC算法结合并利用熵正则化策略优化,经SAR-2数据集训练评估,HRL-SAC在多方面优于普通SAC,为相关任务提供有前途的解决方案。
AI 中文摘要
在稀疏奖励长视野任务中的探索对强化学习提出了重大挑战。为应对这些挑战,我们提出了一个两级分层强化学习(HRL)框架。第一级处理高级战略规划,而低级使用连续控制软演员-评论家(SAC)算法,它们利用熵正则化策略优化。所提出的框架使用搜索与救援-2(SAR-2)数据集进行训练和评估。HRL-SAC有效解决了以延迟奖励和连续控制为特征的稀疏奖励长视野搜索问题,并且在成功率、覆盖效率和收敛方面优于普通SAC基线强化学习。这些发现表明分层熵正则化策略是解决长视野稀疏奖励强化学习任务的一个有前途的解决方案。
英文摘要
Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To address these challenges, we propose a two-level Hierarchical Reinforcement Learning (HRL) framework. The first level handles high-level strategic planning, while the low-level uses the continuous-control Soft Actor-Critic (SAC) algorithm, and they utilize entropy-regularized policy optimization. The proposed framework was trained and evaluated using the Search-and-Rescue-2 (SAR-2) dataset. HRL-SAC effectively addresses sparse-reward long-horizon search problems characterized by delayed rewards and continuous control, and its outperforming the flat SAC baseline reinforcement learning in terms of success rates, coverage efficiency, and convergence. These findings indicate that hierarchical entropy-regularized policies are a promising solution to tackle long-horizon sparse-reward reinforcement learning tasks.
Comments27 pages, 6 figures, 11 Tables