AI 中文总结
针对视觉连续控制的样本高效策略学习难题,本文提出OG-SPR算法,结合多步隐自预测与下一观测预测,通过轻量适配器优化表示,在DeepMind Control Suite的28项任务中优于现有最优方法。
AI 中文摘要
从像素中进行样本高效的策略学习是强化学习(RL)中一个长期存在的挑战。近期基于动力学的表示学习方法通过在隐空间(自预测)或观测空间(观测预测)中执行辅助预测,学习感知动力学的表示,显著提升了无模型视觉强化学习的样本效率。然而,来自这两类方法的当前最优方法在训练数据有限的情况下,仍难以应对具有挑战性的视觉控制任务。我们认为仅依赖单一预测目标可能不够:观测预测将学习到的表示建立在观测级动力学的基础上,但未直接正则化隐表示在扩展时域上的可预测性。本文中,我们提出基于观测的自预测表示(OG-SPR),这是一种用于连续控制的无模型视觉强化学习算法,其学习到的表示既在隐空间中具有时域可预测性,又建立在观测级动力学的基础上。OG-SPR包含两个核心辅助目标:多步隐自预测和下一观测预测。我们通过实验表明,直接对共享表示施加隐自预测可能会过度约束它,且不一定能提升性能。为解决该问题,OG-SPR为隐自预测引入了两个轻量适配器,使共享表示能从时域预测信号中获益,而无需被迫直接满足自预测目标。在来自DeepMind Control Suite的28个视觉控制任务上进行的实验显示,OG-SPR相比当前最优的自预测和观测预测强化学习方法,提升了整体性能,在狗、类人机器人等具有挑战性的领域中增益尤为显著。
英文摘要
Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly improved the sample efficiency of model-free visual RL by learning dynamics-aware representations through auxiliary prediction performed either in latent space (self-prediction) or observation space (observation prediction). However, state-of-the-art methods from both categories still struggle on challenging visual control tasks when training data is limited. We posit that relying on either predictive objective alone may be insufficient. In contrast, observation prediction grounds learned representations in observation-level dynamics, but does not directly regularize the temporal predictability of latent representations over extended horizons. In this paper, we propose Observation-Grounded Self-Predictive Representations (OG-SPR), a model-free visual RL algorithm for continuous control that learns representations that are both temporally predictive in latent space and grounded in observation-level dynamics. OG-SPR incorporates two core auxiliary objectives: multi-step latent self-prediction and next-observation prediction. We empirically show that directly imposing latent self-prediction on the shared representation may over-constrain it and does not necessarily improve performance. To address this issue, OG-SPR introduces two lightweight adapters for latent self-prediction, allowing the shared representation to benefit from temporally predictive signals without being forced to directly satisfy the self-prediction objective. Experiments on 28 visual control tasks from the DeepMind Control Suite show that OG-SPR improves aggregate performance over state-of-the-art self-predictive and observation-predictive RL methods, with particularly pronounced gains in challenging domains such as dog and humanoid.