发表机构
Aalto University; Tampere University(阿尔托大学; 坦佩雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出可达性感知的目标选择框架SUN,结合新颖性与可达性,可集成于离策略RL算法,并在多种复杂环境中超越现有方法。
AI 中文摘要
强化学习(RL)中的探索仍然是一个基本挑战。最近的目标条件强化学习策略(通过选择目标以鼓励更广泛的状态覆盖)已显示出有前景的结果,但没有任何方法能同时以新颖性和可达性为目标:这两个信号要么被手动权衡,要么按顺序应用,要么其中一个被完全忽视。在本文中,我们引入了一个可达性感知的目标选择框架,该框架明确整合了这两个方面,并且可以无缝地集成到任何离策略强化学习算法中。为此,我们提出了后继到新颖性(SUN),这是一个从后继价值函数导出的指标,用于识别既新颖又可达的目标。我们证明了SUN在极限情况下恢复了基于计数的奖励,限制了短视界命中概率,并且可证明地拒绝了不可达的目标。我们进一步提出了一种利用这些属性的自适应目标选择策略,以及一种准确且轻量级的伪计数,以避免经典方法的开销。我们通过全面的基准测试支持我们的所有主张:在具有不可达或难以到达状态、不可逆转换、障碍、迷宫和无界空间的标准和新颖环境中,SUN始终优于最先进的方法。
英文摘要
Exploration in reinforcement learning (RL) remains a fundamental challenge. Recent goal-conditioned RL strategies (which select goals to encourage broader state coverage) have shown promising results, but none scores a goal by novelty and reachability jointly: the two signals are traded off by hand, applied in sequence, or one is neglected outright. In this paper, we introduce a reachability-aware goal-selection framework that explicitly integrates these two aspects, and that can be seamlessly incorporated into any off-policy RL algorithm. To this aim, we propose SUccessor-to-Novelty (SUN), an indicator derived from successor value functions to identify goals that are both novel and reachable. We prove that SUN recovers count-based bonuses in the limit, bounds short-horizon hitting probabilities, and provably rejects unreachable goals. We further present an adaptive goal-selection strategy that leverages these properties, and an accurate yet lightweight pseudocount to avoid the overhead of classic methods. We back up all our claims with thorough benchmarks: SUN consistently outperforms state-of-the-art methods in standard and novel environments with unreachable or hard-to-reach states, irreversible transitions, obstacles, mazes, and unbounded spaces.
Comments39 pages. Accepted at the 19th European Workshop on Reinforcement Learning (EWRL 2026)