arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13035cs.LGcs.AI

基于群胚的内部状态表示用于具有局部对称性的强化学习

Groupoid-Based Internal State Representations for Reinforcement Learning with Local Symmetries

Ben Opperman, Eduardo Alonso, Esther Mondragón

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有强化学习难以利用局部对称性的问题,提出基于群胚的框架,动态发现等价结构,在缩减空间学习,提升样本效率和收敛性,优于标准Q学习。

中文摘要 AI 辅助

对称性在降低强化学习问题的复杂性方面发挥着核心作用,然而大多数现有方法依赖于固定的群作用或预定义的状态抽象。经典强化学习算法通常假设一个全局结构化的马尔可夫决策过程,具有统一适用的动作和转移,这一假设限制了它们利用许多现实环境中存在的模块化以及局部、上下文相关的规律性的能力。我们提出了一种使用群胚来捕获局部、状态依赖的对称性,并在交互过程中支持等价结构的动态发现的强化学习框架。智能体维护轨道代表以及将原始状态映射到规范形式的传输器,使得学习和决策能够在对称性缩减的空间中执行,同时保留局部差异。实证结果表明,所提出的基于群胚的方法在具有强部分对称性的密集和大规模环境中提高了样本效率和收敛性,相较于标准Q学习产生了显著的性能提升。这些发现表明,动态利用局部对称性为可扩展和可泛化的强化学习提供了一条实用且数学上原则性的途径。

英文摘要

Symmetries play a central role in reducing the complexity of reinforcement learning problems, yet most existing approaches rely on fixed group actions or predefined state abstractions. Classical reinforcement learning algorithms typically assume a globally structured Markov decision process with uniformly applicable actions and transitions, an assumption that limits their ability to exploit modularity and local, context-dependent regularities present in many realistic environments. We propose a reinforcement learning framework using groupoids to capture local, state-dependent symmetries and support the dy- namic discovery of equivalence structures during interaction. The agent maintains orbit representatives together with transporters that map raw states to canonical forms, enabling learning and decision-making to be performed in a symmetry-reduced space while preserving local distinctions. Empirical results demonstrate that the proposed groupoid-based approach improves sample efficiency and convergence in dense and large-scale environments exhibiting strong partial symmetries, yielding substantial performance gains over standard Q-learning. These findings show that dynamically exploiting local symmetry provides a practical and mathematically principled route to scalable and generalisable reinforcement learning.

发表机构

  • City St George’s, University of London(伦敦城市圣乔治大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑