Scale-free adaptive planning for deterministic dynamics & discounted rewards
无标度自适应规划用于确定性动力学与折扣奖励
机构 * University of California, Berkeley, USA(加州大学伯克利分校) ; Noah's Ark Lab, Huawei Technologies, London, UK(华为技术伦敦诺亚实验室) ; Adobe Research, San Jose, USA(Adobe研究实验室) ; SequeL team, INRIA Lille - Nord Europe, France(INRIA里尔-北欧洲SequeL团队)
AI总结 本文提出Platypoos算法,针对确定性动力学和折扣奖励的规划问题,提供改进的样本复杂度分析,并建立下界证明其最优性。
Comments 36th International Conference on Machine Learning (ICML 2019)
Journal ref Proceedings of the 36th International Conference on Machine Learning (ICML 2019)