arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39901cs.LGcs.AI

更好的目标表示能否提升目标条件强化学习?

Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?

  • University of Dhaka(达卡大学)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Syed Nazmus Sakib, Abdul Monaf Chowdhury, Nafiul Haque, Shifat E Arman, Md Mehedi Hasan

中文总结 AI 辅助

本研究通过离线迷宫实验发现,目标表示质量对性能影响甚微,而当前状态表示才是关键瓶颈,并提出随机傅里叶位置编码显著提升困难导航任务性能。

中文摘要 AI 辅助

目标条件强化学习(GCRL)在很大程度上依赖于目标如何被表示给策略。尽管近期方法通过时间距离、占用度或可控性来编码目标,但下游性能实际上在多大程度上依赖于表示质量仍不清楚。我们在离线GCRL中通过构建确定性迷宫中的精确时间距离目标表示来研究这一问题。然后,我们在保持下游学习者不变的情况下,系统地破坏其几何质量。在OGBench导航任务和两种算法中,目标表示质量的大幅变化几乎不引起性能变化。然而,对智能体当前状态施加相同的干预,成功率提升超过两倍,揭示了状态路径才是真正的瓶颈。基于这一洞察,我们展示了简单的随机傅里叶位置编码在无需地图信息或目标修改的情况下,显著提升了最困难导航任务的性能。总体而言,我们的发现表明,在基于状态的离线导航中,改进智能体当前状态的表示远比优化目标表示更为重要。代码即将发布。

英文摘要

Goal-conditioned reinforcement learning (GCRL) relies heavily on how target goals are represented to the policy. While recent methods encode goals via temporal distance, occupancy, or controllability, it remains unclear how much downstream performance actually depends on representation quality. We study this in offline GCRL by constructing an exact temporal-distance goal representation in deterministic mazes. We then systematically corrupt its geometric quality while keeping the downstream learner fixed. Across OGBench navigation tasks and two algorithms, large changes in goal-representation quality produce almost no change in performance. However, applying the same interventions to the agent's current state more than doubles success, revealing the state pathway as the true bottleneck. Building on this insight, we show that simple random Fourier positional encodings substantially improve performance on the hardest navigation tasks without map information or objective modifications. Overall, our findings suggest that in state-based offline navigation, improving how the agent's current state is represented matters far more than refining the goal representation. Code will be released soon.

补充信息

↑