AI 中文总结
受人类学习的重复机制启发,提出即时回合重复(IER)机制,将其集成至 SAC、TD3 算法并在多类基准任务上验证,可提升强化学习的样本效率与学习性能。
AI 中文摘要
重复是人类学习的基本机制,重访成功经历可强化记忆、巩固技能并提升未来表现。受这一生物学原理启发,我们提出即时回合重复(Instant Episode Repetition, IER)这一简单新颖的机制,通过在环境交互中立即重复成功回合的动作序列来提升样本效率。与经验回放、自模仿学习(Self-Imitation Learning, SIL)等在训练更新阶段被动复用过往经验的传统方法不同,IER 直接作用于数据收集过程。当智能体识别到高奖励回合后,会在后续固定数量的回合中重复其动作序列,通过与环境的再次交互强化有价值的行为。我们将 IER 集成至最先进的 SAC 和 TD3 算法中,并在连续控制基准任务(包括 MuJoCo、DeepMind Control Suite 及机械臂的真实动态物体平移任务)上评估其有效性。实验结果表明,这一简单机制相较于标准基线和基于自模仿学习的基线,提升了学习性能。
英文摘要
Repetition is a fundamental mechanism in human learning, where revisiting successful experiences strengthens memory, consolidates skills, and improves future performance. Motivated by this biological principle, we introduce Instant Episode Repetition (IER), a simple and novel mechanism that improves sample efficiency by immediately repeating action sequences from successful episodes during environment interaction. Unlike conventional approaches such as Experience Replay and Self-Imitation Learning (SIL), which passively reuse past experience during training updates, IER directly influences the data collection process. Upon identifying a high-reward episode, the agent repeats its action sequence for a fixed number of subsequent episodes, reinforcing valuable behaviors through renewed interaction with the environment. We integrate IER into state-of-the-art SAC and TD3 algorithms and evaluate its effectiveness on continuous-control benchmarks, including MuJoCo, the DeepMind Control Suite, and a real-world dynamic object translation task with a robotic manipulator. Experimental results demonstrate that this simple mechanism improves learning performance over standard and self-imitation-based baselines.
Comments23 pages, 12 figures. Accepted at RLC 2026; to appear in Reinforcement Learning Journal (RLJ) 2026. Code: https://github.com/UoA-CARES/instant-episode-repetition