发表机构
Regis University(里吉斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究元强化学习中拟蒙特卡罗权重初始化的有效性,通过多种采样方法聚合最优先验,与现代正交默认值对比,发现其在连续控制环境训练收敛上有改进,不同任务中正交方向无偏搜索更优。
AI 中文摘要
本文探讨了拟蒙特卡罗(QMC)权重初始化在现代基准环境中用于元强化学习的有效性。使用各种采样方法来限制基于种群的搜索,并从一组基线任务中聚合最优先验。与现代正交(SB3)默认值相比,当外推到类似的未见连续控制环境时,QMC元先验在训练收敛方面有改进。在不同任务中,正交方向在无偏搜索方面全局更优。
英文摘要
This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning within modern benchmark environments. Various sampling methods are used to bound a population-based search and aggregate an optimal prior from a baseline set of tasks. The QMC meta-priors show improvements in training convergence compared to modern orthogonal (SB3) defaults when extrapolated to similar unseen continuous control environments. In dissimilar tasks, the orthogonal orientation was globally superior for an unbiased search.
Comments6 pages, 6 figures, 1 table