Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning
通过在线强化学习将大语言模型接地于交互环境
机构 * Inria (Flowers)(Inria(Flowers)) ; University of Bordeaux(波尔多大学) ; Hugging Face ; Univ Angers, LERIA, SFR MATHSTIC(昂热大学,LERIA,SFR MATHSTIC) ; Sorbonne Université(索邦大学)
AI总结 本文提出GLAM方法,通过在线强化学习将大语言模型接地于交互环境,以提升样本效率和泛化能力,并探讨在线学习的影响。
Comments The associated code can be found at https://github.com/flowersteam/Grounding_LLMs_with_online_RL. This is an extended version of the paper published at ICML 2023: https://proceedings.mlr.press/v202/carta23a
Journal ref PMLR 202 (2023):3676-3713