等变强化学习的样本复杂度
Sample Complexity of Equivariant Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文证明利用群对称性可显著降低强化学习的样本复杂度,通过理论界和实验验证,对称感知策略学习在高维连续机器人模拟中提升了样本效率与性能。
中文摘要 AI 辅助
强化学习(RL)是机器人控制的一个强大框架,然而其实际应用常受高样本复杂度的阻碍。这在物理领域中尤为受限,因为交互数据成本高昂。尽管世界常展现出几何和物理对称性,标准RL算法通常未能利用这一结构。在本文中,我们证明利用群对称性显著降低了RL的样本复杂度。聚焦于有限时域马尔可夫决策过程,我们发现利用群对称性诱导的同态显著降低了达到最优回报所需环境交互次数的理论上下界。我们进一步将这些界扩展到连续状态和动作空间,在适当的正则性假设下提供相应的样本复杂度保证。在理论之外,我们通过受控实验验证了我们的发现,并展示了在面向高维连续机器人模拟中,对称感知策略学习的优势。我们的结果表明,将对称性整合到学习流程中能在样本效率和性能上带来显著提升,为更数据高效的机器人学提供了一条有原则的路径。
英文摘要
Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.
发表机构
- New Theory AI(新理论人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。