Double Horizon Model-Based Policy Optimization
双视界模型驱动策略优化
机构 * Advanced Telecommunications Research Institute(先进电信研究所) ; Kyoto University(京都大学) ; The University of Tokyo(东京大学)
AI总结 双视界模型驱动策略优化通过分阶段rollout策略,平衡分布偏移、模型偏差和梯度稳定性,在连续控制任务中提升样本效率和运行效率。
Comments Accepted to Transactions on Machine Learning Research (TMLR) Code available at https://github.com/4kubo/erl_lib