arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19850cs.RO

GR2PO:面向连续机器人控制的组相对回报策略优化

GR2PO: Group Relative Return Policy Optimization for Continuous Robot Control

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Pengqin Wang, Qiming Zhang, Shaojie Shen, Jun Ma

AI总结:

GR2PO提出一种无评论家的强化学习框架,通过估计折扣回报和组归一化,在连续机器人控制中超越即时奖励基线,媲美演员-评论家方法,并可在边缘设备部署。

AI中文摘要:

演员-评论家架构已广泛用于连续机器人控制。然而,它们依赖于学习价值网络,在训练过程中引入了额外的计算开销。此外,策略学习也可能受到价值估计近似误差的影响。无评论家的组相对策略优化方法通过移除评论家提供了更简单的训练方式。然而,在密集奖励环境中直接将即时奖励应用于策略优化时,它们无法学习长期行动结果。为解决这些问题,我们提出了组相对回报策略优化(GR2PO),一种用于连续机器人控制的无评论家强化学习框架。GR2PO从并行收集的轨迹中估计折扣回报,在每个 rollout 时间索引处执行组归一化,并使用相对优势和裁剪目标来更新策略。为评估所提出框架的有效性,我们在机器人控制仿真环境中实例化它,并将模型部署到真实世界的边缘设备上。结果表明,GR2PO显著优于使用即时奖励的无评论家基线,并与最先进的演员-评论家方法表现相当。此外,GR2PO展示了具有竞争力的训练效率。在NVIDIA Jetson TX2上的推理测试证明了在边缘平台上部署学习策略的可行性。进一步的消融实验分析了并行组大小、回报估计方法和目标裁剪比率对学习性能的影响。为支持后续研究,论文被接收后我们将公开完整代码,包括框架实现、实验配置以及训练和评估脚本。

英文摘要:

Actor-critic architecture has been widely used in continuous robot control. However, they rely on learning a value network, introducing additional computational overhead during training. Moreover, policy learning may also be affected by the approximation error of value estimation. Critic-free group relative policy optimization methods provide a simpler training approach by removing the need for a critic. However, they fail to learn long-term action outcomes when directly applying immediate rewards to policy optimization in dense-reward environments. To address these problems, we propose Group Relative Return Policy Optimization (GR2PO), a critic-free reinforcement learning framework for continuous robot control. GR2PO estimates the discounted returns from the parallelly collected trajectories, performs group normalization at each rollout time index, and uses relative advantages and clipped targets to update the policy. To evaluate the effectiveness of the proposed framework, we instantiate it on robot control simulation environments and deploy the model to a real-world edge device. The results show that GR2PO significantly outperforms critic-free baselines that use immediate rewards and performs competitively against state-of-the-art actor-critic methods. Furthermore, GR2PO demonstrates competitive training efficiency. Inference tests on NVIDIA Jetson TX2 demonstrate the feasibility of deploying the learned policies on edge platforms. Further ablation experiments analyze the effects of parallel group size, return estimation methods, and target clipping ratio on learning performance. To support follow-up research, we will make the complete code publicly available after the paper is accepted, including the framework implementation, experimental configuration, and training and evaluation scripts.

↑