发表机构
University of Massachusetts, Lowell(马萨诸塞大学洛厄尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究分层强化学习中高层智能体子目标选择问题,提出用粗粒度动力学提供动力学感知内在动机,通过最小化预测不确定性稳定高层策略,经实验验证该方法在非平稳长期环境中表现优于现有方法。
AI 中文摘要
分层强化学习旨在将战略规划与原始执行分开,在解决长期复杂任务中广泛成功,但高层智能体面临稀疏、延迟反馈。本文研究能否通过为高层智能体提供动力学感知内在动机,使其更具策略性地选择子目标。因基于原始转移动力学的动机需广泛覆盖状态 - 动作空间,故提出使用粗粒度动力学,通过学习最小化与粗粒度动力学相关的预测不确定性来稳定高层策略,并提供导航结构。用混合密度网络近似评估不同离散度量来建模预测不确定性。实验表明,密集的、动力学感知内在奖励导致规避风险的子目标选择,使其在非平稳长期环境中优于现有分层强化学习方法。
英文摘要
Hierarchical Reinforcement Learning (HRL) intends to separate strategic planning from primitive execution. It has been widely successful in solving long-horizon and complex tasks, where flat-RL algorithms have difficulty in learning. However, while the low-level agent in HRL benefits from dense feedback and abundant trial opportunities, the high-level agent receives sparse, delayed feedback from the environment and its performance depends on the low-level execution capability. In this paper, we study whether subgoal selection by the high-level agent can be performed more strategically, by providing it with dynamics-aware intrinsic motivation. Since motivation based on primitive transition dynamics would require broad coverage of the state-action space, we propose to use coarse dynamics, i.e., environment transitions aggregated over multiple steps at the temporal scale at which the high-level agent operates. This approach stabilizes the high-level policy by learning to minimize the predictive uncertainty associated with the coarse dynamics, and provides a guided structure for navigation. We model the predictive uncertainty by evaluating different dispersion metrics as approximated by a Mixture Density Network (MDN). Empirically, we observe that a dense, dynamics-aware intrinsic reward leads to risk-averse subgoal selection, enabling it to outperform state-of-the-art HRL methods in non-stationary long-horizon environments.
CommentsManuscript accepted to the Eighteenth Workshop on Adaptive and Learning Agents (ALA), at the 25th International Conference of Autonomous Agents and Multi Agent Systems 2026