发表机构
Nankai University; Peking University; Shanghai Waybot Technology Co., Ltd.; Wuhan University(南开大学; 北京大学; 上海微步科技有限公司; 武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GACA,一种基于不确定性关键性代理的自适应粒度信用分配方法,在长程LLM智能体强化学习中优于GRPO和GiGPO。
AI 中文摘要
强化学习现在是在长程任务上训练大型语言模型智能体的标准方式,在这些任务中,数十个相互依赖的动作之后才有一个稀疏奖励。无评论家的、组相对的方法如GRPO适用于这种场景,但它们将一个轨迹级别的标量广播到每一步,无法说明哪个决策导致了结果。GiGPO通过分组共享锚定状态的时间步来恢复步级信号,但它将步级和回合级估计以固定权重合并,在关键的分支决策和常规的、近乎确定性的转换上花费相同的分辨率。我们认为正确的分辨率是状态相关的,并提出GACA,一种无评论家的估计器,其粒度遵循基于不确定性的关键性代理。GACA通过其自身轨迹已记录的负对数似然对每一步进行评分,然后将两种优势以随该评分增长的每步权重混合,因此梯度在高于平均NLL处放置更多权重于细粒度信号,在低于平均NLL处放置更多权重于回合级信号。我们为实现的混合推导了精确的风险分解,并表明在正方向对齐下,足够小的调制优于固定混合。一个单独的条件结果使用期望NLL界定了局部动作值变化,而误差投影分析刻画了混合何时在标量不确定性重加权之外增加价值。在ALFWorld和WebShop上,GACA在1.5B和7B规模下均优于GRPO和GiGPO的任务成功率。
英文摘要
Long-horizon language-model agents trained with reinforcement learning oftenreceive sparse outcome rewards that do not reveal which decisions along a tra-jectory deserve credit. Episode-level advantages provide coarse trajectory-widecredit, while step-level comparisons offer finer resolution with context-dependentestimation noise. We propose Granularity-Adaptive Credit Assignment (GACA),a critic-free method that adaptively mixes episode- and step-level credit for eachdecision during policy optimization. GACA normalizes the sampled response'smean per-token negative log-likelihood (NLL) within each trajectory and uses theresulting criticality score to determine the step-specific mixture. The computationreuses rollout log-probabilities without additional training rollouts or model eval-uations. Our analysis characterizes optimal score-dependent mixing and derivesconditions linking expected NLL to a lower bound on the preferred step-levelweight. Across ALFWorld and WebShop with 1.5B and 7B backbones, GACAachieves the highest reported mean success rates among the compared methods,while introducing negligible additional computation.
CommentsPreprint