arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有复数(值)记忆的强化学习

Reinforcement Learning with Complex (valued) Memories

Sathya Kamesh Bhethanabhotla, Efstratios Gavves, André Biedenkapp

arXiv 2609.38598首次发表:更新:

发表机构

University of Amsterdam; Karlsruhe Institute of Technology(阿姆斯特丹大学; 卡尔斯鲁厄理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出三种酉循环网络作为循环PPO的替代,利用复数表示的相位信息提升部分可观测环境中的记忆与决策,在rocksample和Craftax等任务中奖励提升2-3倍。

AI 中文摘要

部分可观测环境对深度强化学习构成了根本性挑战,要求智能体从观测中压缩时间信息并维护记忆以做出有效决策。尽管存在从门控循环到注意力机制和基于模型的强化学习等多种方法,但寻找能够捕获长期依赖关系的有效表示技术仍是一个活跃的研究领域。在这项工作中,我们重新审视了酉循环网络(uRNNs)[Arjovsky et al., 2016, Jing et al., 2017],该网络展示了优越的梯度流和联想记忆,将循环和隐藏状态表达在复数向量空间中。其保范的酉动力学使得信息能够通过长序列进行传播。为此,我们提出了三种不同版本的uRNNs作为循环PPO架构的即插即用替代品,并证明简单的循环结构和复数表示相位带来的额外自由度,在多个可改善记忆的任务(包括连续控制)上相对于基线取得了显著提升。我们进一步探索如何通过类比量子态的测量方式,保留复数隐藏状态的相位信息以实现相位感知策略。我们的方法在rocksample和Craftax等环境中相比基线获得了高达2-3倍的奖励,这项工作为强化学习的表示和部分可观测性问题指出了令人兴奋的新方向。代码可在以下网址获取:this https URL

英文摘要

Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a phase-aware policy by drawing a parallel to how quantum states are measured. With our methods reaching up to 2-3 $\times$ the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability. Code is available at: https://github.com/Sathya98/qurl

Comments19 pages, 4 figures, 9 tables Accepted at the 19th European Workshop on Reinforcement Learning 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑