发表机构
Fudan University; CASIA; Shanghai Innovation Institute(复旦大学; 中国科学院自动化研究所; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程任务中奖励稀疏问题,提出反馈丰富环境(FEEs)范式,通过环境端观察丰富替代行动引导,在SciWorld和BFCL基准上显著提升多种RL算法的性能。
AI 中文摘要
大型语言模型在静态推理方面表现出卓越的能力,但通过强化学习(RL)将其训练为自主智能体以完成长时程任务时,常常受到严重的奖励稀疏性的阻碍。虽然通过监督微调(SFT)进行传统的智能体端预热可以缓解这一问题,但这种方法常常受到数据稀缺和探索受限的限制。为了解决这一问题,我们提出了一种范式转变,即通过构建反馈丰富环境(FEEs)来实现环境端适应。通过一项试点研究,我们建立了一种反馈设计策略,该策略通过从行动引导转向观察丰富来重构环境,这一转变发生在情节内探索和情节间演化的后期阶段。在SciWorld和BFCL基准上,使用各种Qwen3模型规模和RL算法(如GRPO、GSPO和DAPO)进行的大规模实验表明,FEEs在标准设置上持续带来性能提升。此外,我们的分析揭示,使用FEEs进行训练(1)通过降低熵波动稳定了训练动态,(2)促进了困难任务中的主动状态空间探索,(3)确保环境引导内化为策略权重,而非仅仅作为推理时的先验,(4)识别出组内反馈一致性作为稳定优化的关键边界。
英文摘要
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional \textit{agent-side warming} up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to \textit{environment-side adaptation} by constructing \textbf{F}eedback-\textbf{E}nriched \textbf{E}nvironments (\textbf{FEEs}). Through a pilot study, we establish a feedback design strategy that reformulates environments by transitioning from action guidance to observation enrichment during the later stages of both intra-episode exploration and inter-episode evolution. Large-scale experiments on SciWorld and BFCL benchmarks using various Qwen3 model scales and RL algorithms such as GRPO, GSPO, and DAPO demonstrate that FEEs consistently yield performance improvements over standard settings. Furthermore, our analysis reveals that training with FEEs \textbf{(1)} stabilizes training dynamics by reducing entropy volatility, \textbf{(2)} facilitates proactive state-space exploration in difficult tasks, \textbf{(3) }ensures the internalization of environmental guidance into policy weights rather than acting as a mere inference-time prior, and \textbf{(4) }identifies intra-group feedback consistency as a critical boundary for stable optimization.
Comments21 Pages, 6 Figures, 7 Tables,