arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

世界模型记忆,智能体遗忘:基于持续模型的强化学习中的梦境排练

The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL

Gurp Nijjer

arXiv 2607.19749首次发表:更新:

发表机构

Quantegra Research(Quantegra研究公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究 DreamerV3 家族智能体在任务序列训练时的遗忘问题,通过组件级探测发现是通道问题,提出分级梦境排练方法,能产生无任务标签、参数恒定的持续学习者,在多任务链保留上表现优异。

AI 中文摘要

DreamerV3 家族的基于模型的强化学习智能体在任务序列训练时会灾难性遗忘,即便有无限回放缓冲区保存所有早期经验。本文提出问题:是哪个组件遗忘?在不清晰回放下,预先注册的组件级探测表明世界模型保留了旧任务的所有可测量信息,而智能体行为崩溃。遗忘是通道问题而非记忆问题,通过干预证明了这一点。训练期间交错进行的分级梦境排练产生了无任务标签、参数恒定的持续学习者,在多任务链保留上优于普通回放,且梦境分级步骤至关重要,本文还刻画了两种评分失败模式并给出应对规则。所有实验均预先注册并报告了被驳斥的假设。

英文摘要

Model-based reinforcement-learning agents of the DreamerV3 family forget catastrophically when trained on task sequences, even when an unbounded replay buffer preserves every earlier experience. We ask a question the continual-RL literature has assumed an answer to but never measured: which component forgets? Under never-clear replay, pre-registered component-level probes (n=3 seeds throughout) show that the world model retains essentially everything measurable about old tasks -- reward discrimination (retention ratio ~1.0), value estimates, and termination structure -- while the actor's behavior collapses. Forgetting in this regime is a channel problem, not a memory problem. We demonstrate this by intervention: with the world model frozen and identical imagined rollouts, reinforcement learning in imagination fails to recover a lost skill (0/3 seeds), while supervised self-imitation on the world model's own graded dreams recovers it on 3/3 seeds with zero environment interaction. Interleaved during training, this graded dream rehearsal yields a task-label-free, parameter-constant continual learner: 3/3 four-task chains retained where plain replay passes 0/3, 3/3 eight-task chains, and consistent gains over matched real-episode cloning (paired difference +0.13, bootstrap 95% CI [0.07, 0.24], complete seed separation). The dream-grading step is load-bearing: we characterize two scoring failure modes, provide an offline selection gauge that caught both before they contaminated results, and give a realized-first grading rule that closes them. All experiments were pre-registered with committed protocols; every refuted hypothesis is reported.

Comments11 pages, 2 figures. Code, pre-registration trail, and run data: https://github.com/gurpnijjer/dream-rehearsal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑