发表机构
Zhejiang University; University of Science and Technology of China; University of New South Wales; Shanghai Artificial Intelligence Laboratory(浙江大学; 中国科学技术大学; 新南威尔士大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出双轴策略优化框架BATON,通过贝叶斯反馈归因和轨迹质量归一化分别优化轨迹内反馈利用与轨迹间聚合,在多个基准上取得一致最优性能。
AI 中文摘要
针对LLM智能体的强化学习涉及两个不同的优化维度:环境反馈如何在轨迹内被利用,以及完整轨迹如何在批次中聚合。我们将这些维度形式化为轨迹内反馈归因和轨迹间目标聚合,并引入BATON(贝叶斯归因与轨迹目标归一化),一种双轴策略优化框架。BATON用贝叶斯反馈归因实例化第一个轴,该归因在采样动作上构建一个以反馈为条件的后验;第二个轴用轨迹质量归一化(TMN)实例化,该归一化赋予完整轨迹相等的优化质量。在ALFWorld、WebShop和SearchQA上使用GRPO和GiGPO进行的实验表明,两个轴都提供了独立的增益,并且它们的组合在不同模型规模上一致地实现了最强的整体性能。
英文摘要
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.