IPD: Boosting Sequential Policy with Imaginary Planning Distillation in Offline Reinforcement Learning
IPD: 通过离线规划蒸馏提升序列策略
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Peking University(北京大学)
专题命中 仿真与规划 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG
AI总结 IPD通过引入离线规划蒸馏技术,提升离线强化学习中序列策略的决策稳定性和性能。