用于强化学习的因式分解谱表示
Factorized Spectral Representations for Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
该研究聚焦强化学习,提出FaStR方法,通过对转移核的三模张量CP分解,用噪声对比目标拟合,产生单独编码器形成谱表示。其因式分解形式缩小假设类,在高维运动任务中效果好,状态编码器可跨执行器移位转移。
中文摘要 AI 辅助
从交互数据中学习世界的紧凑模型是高效样本深度强化学习的核心。谱表示方法通过将转移核视为矩阵,以状态-动作对和下一个状态为两侧,通过自监督对比目标学习低秩分解,已成为连续控制中表示学习的主导范式。我们进一步拓展这一观点。转移核自然是关于状态、动作和下一个状态的三模张量,CP分解为每个模式给出一个特征图。我们提出FaStR,它通过噪声对比目标拟合这种分解,产生单独的状态、动作和下一个状态编码器,共同形成单个谱表示。因式分解形式产生更小的假设类,表示学习所需的样本大小按与状态和动作维度中较小者成比例的因子缩小。实证上,FaStR在动力学与因式分解结构一致的高维运动任务上取得最大收益,并且学习到的状态编码器在执行器移位时完整转移,只需重新训练动作编码器。
英文摘要
Learning a compact model of the world from interaction data is central to sample-efficient deep reinforcement learning. Spectral representation methods have become the leading paradigm for representation learning in continuous control by taking a matrix view of the transition kernel, with state-action pairs on one side and next states on the other, and learning a low-rank factorization through self-supervised contrastive objectives. We take this view one step further. The transition kernel is naturally a three-mode tensor over states, actions, and next states, and a CP decomposition gives one feature map per mode. We propose FaStR, which fits this decomposition with a noise contrastive objective, producing separate state, action, and next-state encoders that together form a single spectral representation. The factored form yields a smaller hypothesis class, and the sample size needed for representation learning shrinks by a factor that scales with the smaller of the state and action dimensions. Empirically, FaStR delivers its largest gains on high-dimensional locomotion tasks whose dynamics align with the factored structure, and the learned state encoder transfers intact across actuator shift while only the action encoder is retrained.
发表机构
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。