arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ChronoSRL:用于自监督强化学习的时序几何

ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning

Nico Bohlinger, Jan Peters

arXiv 2609.36238首次发表:更新:

发表机构

Technical University of Darmstadt; Robotics Institute Germany (RIG); German Research Center for AI (DFKI)(达姆施塔特工业大学; 德国机器人研究所; 德国人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ChronoSRL通过为批评者嵌入赋予时序几何,将状态-动作与目标间的距离训练为匹配到达时间,并预测到达时间分布及目标附近停留时间,从而在多个基准和四足机器人任务上实现更快、更可靠的学习,超越现有自监督强化学习方法。

AI 中文摘要

在空间中距离相近的目标,在时间上可能相距甚远。障碍物、地形以及智能体自身的能力决定了到达目标所需的时间。然而,对比强化学习和生存强化学习中的批评者并未以其表示空间中的距离来衡量时间单位。因此,我们引入了ChronoSRL,它为批评者的嵌入赋予了明确的时序几何。状态-动作嵌入与目标嵌入之间的距离被训练为匹配智能体到达目标所需的时间(到达目标时间),而未被到达的目标以及其他轨迹中的目标则被推离至少一个折扣视界。此外,一次快速到达目标并不意味着通常能可靠地到达目标,因此策略不应直接遵循时序距离。相反,我们基于生存强化学习,从我们的时序嵌入中不仅预测到达目标时间的完整分布,还预测在目标附近停留的时间。由此,策略被训练为倾向于选择能更快、更可靠地到达目标并保持智能体靠近目标的动作。在七个标准运动与导航基准上,ChronoSRL比对比强化学习、动作分块对比强化学习和生存强化学习基线学习得更快,并达到更高的性能,即使使用更小的网络也是如此。为了测试自监督强化学习的极限,我们在一个逼真的仿真到现实运动设置中,使用四足机器人引入了速度跟踪、目标位置到达和箱子攀爬任务,并展示了机器人技术中典型的塑形项如何自然地融入我们的框架。ChronoSRL是所测试的自监督强化学习方法中唯一学会保持在指令速度和目标位置并攀爬最高箱子的方法。

英文摘要

A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their representation space in units of time. We therefore introduce ChronoSRL, which gives the critic's embeddings an explicit temporal geometry. The distance between state-action and goal embeddings is trained to match the time that the agent takes to reach the goal (goal-reaching time), while goals that were not reached, and goals from other trajectories, are pushed at least one discount horizon away. Furthermore, reaching a goal quickly once does not mean that reaching it is reliable in general, so the policy should not follow the temporal distance directly. Instead, we build on survival reinforcement learning and predict from our temporal embeddings not only the full distribution of goal-reaching times but also the time spent near the goal. Thereby, the policy is trained to favor actions that reach the goal sooner and more reliably and that keep the agent near it. ChronoSRL learns faster and reaches higher performance than contrastive, action-chunked contrastive, and survival reinforcement learning baselines on seven standard locomotion and navigation benchmarks, even with much smaller networks. To test the limits of self-supervised reinforcement learning, we introduce velocity tracking, goal-position reaching, and box climbing tasks with a quadruped robot in a realistic sim-to-real locomotion setup, and show how the shaping terms that are typical for robotics can be naturally incorporated into our framework. ChronoSRL is the only one of the tested self-supervised reinforcement learning methods that learns to stay at the commanded velocities and goal positions, and climbs the highest boxes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑