Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
机构 * The University of Tokyo(东京大学)
专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.AI、cs.LG
Comments Accepted at ICML 2025. Source code: https://github.com/motokiomura/annealed-q-learning