Internalizing World Models via Self-Play Finetuning for Agentic RL
机构 * City University of Hong Kong(香港城市大学) ; Northwestern University(西北大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; Oxford University(牛津大学) ; Allen Institute for AI (AI2)(人工智能研究所) ; University of Washington(华盛顿大学) ; National University of Singapore(新加坡国立大学) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 模仿学习与强化学习 :world model(title,abstract);分类 cs.LG