arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ALER:用于强化学习的自适应可学习经验重写

ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning

Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev

arXiv 2610.00592首次发表:更新:

发表机构

Innopolis University; MIRIAI; Cognitive AI Systems Lab(英诺波利斯大学; MIRIAI; 认知AI系统实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对部分可观测强化学习中信息过时问题,提出ALER智能体,结合LSTM与槽记忆,通过可学习重写和融合机制,在多个基准上显著提升成功率。

AI 中文摘要

在部分可观测的强化学习(RL)中,后来的观测可能使存储的信息过时,或改变其对下一步决策的含义。用于RL的记忆架构和基准大多测试保留能力,即保持信息不变直到需要它的能力。我们形式化了另外两个要求。重写将决策相关内容设置为独立于旧内容的值,而经验融合通过后续观测指定的规则转换旧内容。对于由这类更新构建的任务,我们计算解决方案所需的记忆状态数量,几个基线在需要更多状态的任务组合上达到了最低成功率。我们引入了ALER(自适应可学习经验重写),一种将LSTM与槽记忆配对的智能体。一个独立寻址的Gumbel-Softmax写入,将其权重集中在一个槽上,覆盖该槽,并且一个学习的门在策略和价值头之前将检索到的内容与循环状态融合。我们还引入了Rune-Mazes,三个环境,其中符文观测在向量和像素观测下反转、取消、重置或重复隐藏线索的更新。与七个基线相比,ALER在所有十六个Endless T-Maze配置中达到了至少0.82的成功率,在所有五个Rune T-Maze组合中达到了至少0.99的成功率,并且在带有Invert符文的四分支Rune Multi-Corridor中具有最高的平均成功率。在基于像素的Rune MiniGrid Memory上,它在十个配置中的八个中比PPO-LSTM具有更高的平均成功率。项目页面:此https URL。

英文摘要

In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least $0.82$ in all sixteen Endless T-Maze configurations and at least $0.99$ on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: https://quartz-admirer.github.io/ALER-Adaptive-Learnable-Experience-Rewriting/.

Comments28 pages, 12 figures, 18 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑