SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling
SAC-GLAM:通过软演员-评论员和回顾重标记改进LLM代理的在线强化学习
机构 * Inria (Flowers)(Inria(Flowers)) ; University of Bordeaux(波尔多大学) ; Hugging Face ; Univ Angers(昂热大学) ; LERIA ; SFR MATHSTIC(MATHSTIC联合研究机构) ; Sorbonne Université(索邦大学) ; ISIR
AI总结 SAC-GLAM通过结合软演员-评论员算法和回顾重标记,改进LLM代理的在线强化学习,提升其在复杂环境中的策略学习能力。
Comments This work has been presented at the IMOL workshop at NeurIPS 2025 (https://neurips.cc/virtual/2024/101058)