从环境反馈中学习:面向智能体强化学习的跨时间尺度信用分配
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对智能体强化学习的奖励延迟稀疏问题,提出基于环境反馈的跨时间尺度信用分配方法EFCA,在ALFWorld和WebShop上提升了任务成功率与质量。
中文摘要 AI 辅助
智能体强化学习(RL)在真实环境中常面临奖励延迟且稀疏的问题,信用分配是解决该挑战的可行方案,其目标是分解轨迹级奖励,为中间决策提供更细粒度的监督。但现有信用分配方法忽略了环境交互中自然产生的丰富过程信息,如交互历史,而这些信息对识别单个动作的贡献具有重要监督价值。为此,我们提出基于环境反馈的信用分配(EFCA),这是一种面向长周期智能体RL的跨时间尺度信用分配方法。EFCA用两个基于环境的过程信号补充长期结果信号:一是捕捉当前动作即时效果的短期反馈信号,二是从近期交互中识别无效模式的中期状态历史信号。两种信号均直接从环境反馈中提取,并通过回报重加权机制整合。在ALFWorld和WebShop上的实验表明,EFCA相较于强基线始终提升了任务成功率和任务质量,凸显了基于环境的跨时间尺度信用分配在长周期智能体RL中的有效性。
英文摘要
Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide more fine-grained supervision for intermediate decisions. However, existing credit assignment approaches ignore the rich process information naturally generated during environment interaction, e.g., interaction history. We argue that such information provides valuable supervision for identifying the contribution of individual actions. To this end, we propose Environmental Feedback-based Credit Assignment (EFCA), a multi-timescale credit assignment approach for long-horizon agentic RL. EFCA complements the long-term outcome signal with two environment-grounded process signals: a short-term feedback signal that captures the immediate effect of the current action and a medium-term state-history signal that identifies ineffective patterns from recent interactions. Both signals are directly extracted from environment feedback and integrated through a return reweighting mechanism. Experiments on ALFWorld and WebShop demonstrate that EFCA consistently improves both task success and task quality over strong baselines, highlighting the effectiveness of environment-grounded multi-timescale credit assignment for long-horizon agentic RL.