端点回放:压缩深度强化学习中的近期缓冲区
Endpoint Replay: Compressing the Recency Buffer in Deep Reinforcement Learning
AI总结:
研究在深度强化学习中压缩近期缓冲区的方法,提出存储相连n步序列链端点派生的代表性转换的压缩方法,经实证评估,该方法防止系统偏差,在相关环境和基准测试中与传统大型缓冲区性能相当。
AI中文摘要:
经验回放仍是深度强化学习(DRL)工具箱中最实用且有用的算法工具之一。除了优先回放以及针对大型异步系统的专门方法取得有限成功外,大多数DRL算法使用大型、均匀采样的近期缓冲区,即便大小达一百万也不变。本文研究一种简单压缩方法,存储从相连的n步序列链端点派生的代表性转换。通过在较小近期缓冲区中精心挑选这些端点,该方法保持与标准大型缓冲区相当的有效记忆跨度,同时存储需求减少一个数量级。实证评估表明,此方法防止了朴素压缩策略中固有的系统偏差,在弹珠台环境和雅达利2600基准测试中与传统大型缓冲区性能相当。
英文摘要:
Experience replay remains one of the most practical and useful algorithmic tools in the deep reinforcement learning (DRL) toolbox. Aside from the limited success of prioritized replay and specialized approaches for large asynchronous systems, most DRL algorithms make use of a large, uniformly sampled recency buffer---even the size, one million, remains unchanged. Could we store less data, reduce redundancy, or more effectively chain experience together to speed up value propagation and still retain the performance of large buffers? In this paper, we investigate a simple compression approach that stores representative transitions derived from the end-points of a chain of connected $n$-step sequences. By curating these end-points in a smaller recency buffer, our method maintains an effective memory horizon comparable to a standard large buffer while requiring an order of magnitude less storage. Through empirical evaluation, we demonstrate that this approach prevents the systematic bias inherent in naive compression strategies and matches the performance of traditional large buffers in the Pinball environment and the Atari 2600 benchmark.