发表机构
Loughborough University; University of Pisa(拉夫堡大学; 比萨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出AMSC方法,利用Wasserstein任务嵌入估计相似性并组合先前策略,在终身强化学习中实现高效迁移,实验显示更高性能且无遗忘。
AI 中文摘要
在终身强化学习中,仅保留先前学到的策略不足以有效迁移到新任务。有用的知识可能分布在多个先前的策略中,并且其相关性会随着学习者的经验积累而变化。一个假设是,在持续学习环境中,任务相似性可被有效用于寻找并组合先前学到的策略。为验证该假设,设计了自适应掩码选择与组合(AMSC)方法,该方法通过基于状态-动作-奖励样本的非参数Wasserstein任务嵌入,从在线经验中估计任务相似性。利用相似度分数的z-score归一化sparsemax,推导出可变大小的支持集,以定期选择并加权策略,从而在学习新任务时形成先验。在CT-graph和MiniGrid上,AMSC相比所评估的模块化组合基线,实现了更高的平均性能和前向迁移,且无遗忘现象。在Continual World上的结果表明,识别相关的先验知识并确定其逐层组合可能需要额外的逐层调优。消融实验显示,选择相关来源并确定复用强度是这些收益的核心。独立测量的成对迁移与任务嵌入相似性也呈正相关。这些结果表明,任务相似性可成为在终身强化学习中选择并加权特定知识以进行复用的有效标准。
英文摘要
In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity can be effectively used in a continual learning setting to find and combine previously learned policies. To test it, Adaptive Mask Selection and Composition (AMSC) is designed to estimate similarity from online experience via non-parametric Wasserstein task embeddings from state-action-reward samples. The z-score-normalized sparsemax of the similarity scores are used to derive a variable-size support to periodically choose and weight policies to form a prior when learning a new task. On CT-graph and MiniGrid, AMSC achieves higher mean performance and forward transfer than the evaluated modular composition baselines while exhibiting no forgetting. Results on Continual World suggest that identifying relevant prior knowledge and determining its layer-specific composition may require additional layer-specific tuning. Ablations show that selecting relevant sources and determining how strongly to reuse them are central to these gains. Independently measured pairwise transfer is also positively associated with task-embedding similarity. These results indicate that task similarity can be an effective criterion to select and weight specific knowledge for reuse in lifelong reinforcement learning.
CommentsCode is available at https://github.com/Chocological45/amsc