GraphHCA:面向长程LLM智能体的闭式后见信用分配
GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents
浏览论文内容
中文总结 AI 辅助
针对长程LLM智能体的稀疏奖励信用分配问题,提出GraphHCA,通过贝叶斯规则和转移图上的折扣递归实现无模型后见信用分配,在ALFWorld、WebShop和Sokoban上取得最先进结果。
中文摘要 AI 辅助
基于组的强化学习(RL)已推动大型语言模型(LLMs)的发展,并日益扩展到智能体任务,在这些任务中,稀疏的终端奖励使得步骤级信用分配至关重要。现有方法根据采样轨迹中动作之后发生的事件来分配信用,但并未明确捕捉其与已实现结果之间的回顾性关系。后见信用分配(HCA)则通过后见概率与行为策略概率之比来分配信用,但估计后见分布需要辅助模型或额外的一次前向传播。为解决这一估计瓶颈,我们提出GraphHCA,一种无模型的HCA实现,消除了显式的后见分布估计。对于具有确定性转移的终端目标任务,贝叶斯规则将后见比率简化为连续状态上行为策略成功概率之比。取对数得到状态级成功势,其在转移中的增量提供步骤级信用。GraphHCA通过诱导转移图上的折扣递归从合并轨迹中估计该势,该递归在任何有向图上均存在唯一不动点。所得步骤级信号与轨迹级优势相结合,既不需要学习后见模型,也不需要额外的前向传播,当步骤级权重为零时恢复GRPO。在所有比较的基线中,GraphHCA在ALFWorld和WebShop上,在两种LLM规模下均取得了最先进的结果,并在Sokoban上使用视觉语言智能体取得了最先进结果。例如,在ALFWorld上,其整体成功率比GRPO提高最多24.6个百分点,比最强的步骤级基线提高最多4.7个百分点。
英文摘要
Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its retrospective relation to the realized outcome. Hindsight credit assignment (HCA) instead attributes credit through the ratio of hindsight to behavior-policy probabilities, but estimating the hindsight distribution requires an auxiliary model or an extra pass. To address this estimation bottleneck, we propose GraphHCA, a model-free realization of HCA that eliminates explicit hindsight-distribution estimation. For terminal-goal tasks with deterministic transitions, Bayes' rule reduces the hindsight ratio to a ratio of behavior-policy success probabilities at consecutive states. Taking logs yields a state-wise success potential, whose increment across a transition provides step-level credit. GraphHCA estimates this potential from pooled rollouts through a discounted recursion on the induced transition graph, which admits a unique fixed point on any directed graph. The resulting step-level signal is combined with the trajectory-level advantage, requiring neither a learned hindsight model nor an extra forward pass and recovering GRPO when the step-level weight is zero. Among all compared baselines, GraphHCA achieves state-of-the-art results on ALFWorld and WebShop at both LLM scales, and on Sokoban with a vision-language agent. For example, on ALFWorld it improves overall success rate by up to 24.6 points over GRPO and by up to 4.7 points over the strongest step-level baseline.
发表机构
- Beihang University(北京航空航天大学)
- Zhongguancun Academy(中关村学院)
- Communication University of China(中国传媒大学)
- Hangzhou Innovation Institute of Beihang University(北京航空航天大学杭州创新研究院)
机构由 AI 辅助整理,请以论文原文为准。