AI 中文总结
研究针对基于存档探索方法未考虑边界后区域能否进入的问题,提出TopoExplore,通过周期性拓扑遍历增强单元格选择,在MiniGrid套件等实验中取得加速效果,证明拓扑感知选择在区分封闭结构及开放覆盖区域均有优势。
AI 中文摘要
基于存档的探索方法如Go-Explore通过访问稀有性选择返回的访问状态,前沿方法返回未知区域边界,但都未考虑边界后的未探索区域是否可进入。探索不仅是为了找到奖励,还需为下游学习和规划收集结构完整的经验。我们引入TopoExplore,通过周期性拓扑遍历增强Go-Explore的单元格选择。利用泛洪填充检测已访问集占用网格中的封闭未探索区域(空洞),仅在其严格入口(间隙或门单元格)放置衰减选择奖励,使密封区域不被目标且已进入区域退出。在18个环境的MiniGrid套件上,TopoExplore在首次进入的中位数步数上比精确的Go-Explore消融实现了1.52倍的几何平均加速,优于前沿基线的1.37倍。前沿探索在出现密封诱饵结构时会退化,而TopoExplore在硬多交互门环境中优势最大(10.9倍)。在蒙特祖玛的复仇游戏中,若无墙壁知识,无法到达的占用工件会捕获奖励并导致性能下降,而在HM3D扫描建筑中,相对于Go-Explore的加速与场景难度相关(r = 0.69)。证据支持一个有针对性的观点:拓扑感知选择在必须区分封闭结构的地方有回报,在开放覆盖区域也具有竞争力,尽管未针对该模式进行调整。
英文摘要
Exploration signals in reinforcement learning are currently computed from what an agent has seen: visitation counts, density estimates, or a model's prediction error at individual states. None of these report the holes in the visitation space. To capture them, we compute the persistent homology of the agent's own archive of visited states during exploration and turn each detected class (an enclosed region the archive surrounds but has not entered) into a selection bonus concentrated on the archived states from which entry is possible, gated by an attempt counter that retires candidate entrances and prunes sealed structures. The method, TopoExplore, is Go-Explore plus one additive term, so the comparison with Go-Explore isolates that term. On the 189 held-out worlds of the open-source TopoGym benchmark of topologically varied hard-exploration environments, TopoExplore finds the goal in 167 worlds against 152 for Go-Explore and, where both find it, with a median 126k steps against 186k. On an environment that stress-tests far-apart chambers that create holes in the archive space, TopoExplore enters all six or eight chambers across 5 seeds, against 1 and 2 of 5 seeds for Go-Explore. On Montezuma's Revenge, built over the room graph, the term helps whenever the archive surrounds a room it has not entered, puts 33-45% of the selections on the surrounding rooms while it is active, and enters the room sooner in five of six cases; such enclosures are rare on this game, so the score is unchanged. Beyond discrete, restorable state spaces, the same signal drives an embodied explorer that walks to every target it selects and explores from depth alone; a vision-language-action navigation policy fine-tuned on its trajectories succeeds 58% more often on unseen scenes than one fine-tuned on shortest-path oracle demonstrations.