arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强化信息梦行者:通过潜在引导有效训练的非对称世界模型

Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance

Gaspard Lambrechts, Adrien Bolland, Daniel Ebi, Damien Ernst

arXiv 2607.26040首次发表:更新:

AI 中文总结

研究基于模型的强化学习中,非对称学习对观测及特权信息表示的影响。针对‘信息梦行者’局限性,提出用潜在引导的新目标,形成‘强化信息梦行者’算法,实验显示其比之前非对称方法有更持续改进。

AI 中文摘要

与人类学习时受益于引导类似,强化学习算法可能从奖励之外的额外监督中受益。在训练期间利用额外信息以学习更好的表示和行为一直是非对称强化学习的重点。这种学习范式在部分可观测性且有额外状态信息时,以及完全可观测性且有更精细状态信息时都已证明有效。聚焦基于模型的强化学习,我们研究非对称学习对观测表示和特权信息表示的影响。首先,我们识别出一种基于模型的非对称算法‘信息梦行者’所学习的特权信息表示中的局限性。然后,我们提出一种使用潜在引导的新型非对称表示学习目标,从而产生一种名为‘强化信息梦行者’的新算法。在多个基准测试上的实验表明,与之前的非对称方法相比,它对梦行者有更持续的改进。

英文摘要

Much like humans benefit from guidance while learning, reinforcement learning algorithms may benefit from additional supervision beyond rewards. Leveraging additional information during training to learn better representations and behaviors has been the focus of asymmetric reinforcement learning. This learning paradigm has proven effective under partial observability when additional state information is available, but also under full observability when more refined state information is available. Focusing on model-based reinforcement learning, we study the effect of asymmetric learning on observation representations and on privileged information representations. First, we identify a limitation in the privileged information representations learned by an asymmetric model-based algorithm known as the Informed Dreamer. Then, we propose a novel asymmetric representation learning objective using latent guidance, resulting in a new algorithm called the Reinformed Dreamer. Experiments across several benchmarks show a more consistent improvement over Dreamer than previous asymmetric approaches.

Comments8 pages, 18 pages total, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑