发表机构
Alibaba Group; Kyoto University; NII LLMC; Peking University; University of California, Los Angeles; Arizona State University; The Chinese University of Hong Kong, Shenzhen; Tsinghua University; University of the Chinese Academy of Sciences(阿里巴巴集团; 京都大学; 国立情报学研究所LLMC; 北京大学; 加州大学洛杉矶分校; 亚利桑那州立大学; 香港中文大学(深圳); 清华大学; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体强化学习中终端奖励导致的信用分配问题,提出FAULT方法,将诊断错误转化为步骤级信用,在ALFWorld上信号覆盖率显著提升。
AI 中文摘要
智能体强化学习已成为训练大型语言模型智能体处理多步任务的有效方法,然而对终端结果奖励的依赖带来了两个信用分配问题,尤其是在长视界任务中。首先,相同结果的轨迹组无法从终端奖励中获得学习信号。其次,终端奖励仅提供轨迹级别的反馈,难以识别哪些决策导致了失败。近期研究通过轨迹分析补充了更细粒度的信息,例如对中间决策和错误的自然语言反思。然而,自然语言诊断难以直接用于信用分配:其错误主张可能不可靠,且未量化每个错误对学习的影响程度。我们提出自我诊断引导的终端信用再分配(FAULT),将诊断出的错误转化为以终端结果锚定的显式步骤级信用。FAULT检查诊断证据并从任务结果中学习相对错误成本。在训练过程中,策略和自诊断器共同进化,而错误成本则根据最近结果在线更新。在ALFWorld上,FAULT从相同结果组中恢复学习信号,达到95%的信号覆盖率,而GRPO为41%,GiGPO为72%,同时更好地将信用定位到特定错误步骤。在两种模型规模下,FAULT在长视界ALFWorld和WebShop任务上取得了显著改进,同时在短视界基于搜索的问答任务上保持竞争力。
英文摘要
Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from terminal rewards. Second, terminal rewards provide only trajectory-wide feedback, making it difficult to identify which decisions caused a failure. Recent work supplements terminal rewards with finer-grained information from trajectory analysis, such as natural-language reflections on intermediate decisions and errors. However, natural-language diagnoses are difficult to use directly for credit assignment: their error claims may be unreliable, and they do not quantify how much each error should affect learning. We propose Self-Diagnosis-guided Terminal Credit Redistribution (FAULT), which turns diagnosed errors into explicit step-level credit anchored by terminal outcomes. FAULT checks diagnostic evidence and learns relative error costs from task outcomes. During training, the policy and self-diagnoser co-evolve, while error costs are updated online from recent outcomes. On ALFWorld, FAULT recovers learning signals from same-outcome groups, reaching 95% signal coverage versus 41% for GRPO and 72% for GiGPO, while better localizing credit to specific error steps. Across two model scales, FAULT delivers strong. improvements on the long-horizon ALFWorld and WebShop tasks while remaining competitive on short-horizon Search-based QA.
CommentsPreprint