arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06861cs.AI

Gated-BEPO:用于大语言模型智能体的置信门控贝尔曼信用分配

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents

Hongxi Yan, Ziyue Huang, Shichao Fan, Qingjie Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对大语言模型智能体长视域训练的信用分配问题,提出Gated-BEPO方法,通过经验rollout图推导步骤级信用并结合置信门自适应融合回合与步骤级信用,在多任务基准上取得性能提升。

中文摘要 AI 辅助

在长视域环境中训练大语言模型智能体,需将稀疏的终端结果信用分配给各个动作。现有无评判者方法会将轨迹级奖励均匀分配到各步骤,而近期方法则通过匹配重复状态构建步骤级分组,并在每组内比较动作。前者无法区分失败轨迹中的有用动作与成功轨迹中的无效动作;后者依赖直接源自单个轨迹结果的步骤信用,以及与回合级信用的固定权重融合。我们提出Gated-BEPO,该方法从经验 rollout 图中推导步骤级信用:对每个 rollout 分组,Gated-BEPO 构建经验图,并通过均值备份贝尔曼不动点估计节点值,该不动点反映当前策略的经验动作分布;随后使用广义优势估计沿每条采样轨迹累积这些时间差残差,生成能同时捕捉即时及下游效应的步骤级贝尔曼优势。为自适应融合回合级与步骤级信用,置信门仅在存在多个观测后继状态时才纳入贝尔曼信用,否则使用回合级信用。在WebShop、ALFWorld及视觉Sokoban上的实验显示,该方法在语言及视觉-语言模型上均取得一致提升,而诊断性 ablation 验证了贝尔曼不动点值估计的有效性,并表明步骤级信用应选择性而非均匀地纳入最终优势计算。

英文摘要

Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual actions. Existing critic-free methods propagate trajectory-level rewards uniformly across steps, while recent approaches construct step-level groups by matching repeated states and compare actions within each group. The former cannot distinguish useful actions in failed trajectories from ineffective actions in successful ones. The latter rely on step credit derived directly from individual trajectory outcomes and fixed-weight fusion with episode-level credit. We propose Gated-BEPO, which derives step-level credit from empirical rollout graphs. For each rollout group, Gated-BEPO constructs an empirical graph and estimates node values through a mean-backup Bellman fixed point that reflects the empirical action distribution of the current policy. We then accumulate these temporal-difference residuals along each sampled trajectory using generalized advantage estimation, yielding step-level Bellman advantages that capture both immediate and downstream effects. To adaptively fuse episode- and step-level credit, a confidence gate incorporates Bellman credit only at states with multiple observed successors and otherwise uses episode-level credit. Experiments on WebShop, ALFWorld, and visual Sokoban show consistent improvements across language and vision-language models, while diagnostic ablations support the effectiveness of Bellman fixed-point value estimation and show that step-level credit should be incorporated selectively rather than uniformly into the final advantage.

↑