arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07118cs.AI

那么多少,然后在哪里:多轮智能体强化学习中保留信用的动作到令牌分配

How Much, Then Where: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning

Lichao Ma, Yang Sun, Shuaitao Zhao, Yangyi Fang, Cong Qin, Xiaoliang Fu, Yuhang Tian, Yuchen Wei, Junbo Zhu, Yang Wei, Lu Pan, Jiaye Lin

AI总结:

本文提出FACTOR算法,通过分离动作级与令牌级信用分配,在多轮智能体强化学习的ALFWorld等环境中实现性能提升,且超参数可跨模型迁移。

AI中文摘要:

多轮智能体强化学习中的信用分配在两个层面进行:将轨迹级信用分配给动作,并将每个动作的信用分配到其令牌上。本文提出了FACTOR,它将这些决策分开。FACTOR使用经检查点校准的TD残差分配每个动作的信用,该信用会扩展到轨迹优势;并使用反馈条件的师生似然差距将每个信用分配到已实现的动作令牌上。每个动作的归一化保留了动作平均系数,防止令牌级符号翻转。我们将此构造与动作平均约简相结合,消除了动作的标量替代权重对其令牌长度的隐式依赖。在行为策略中且在裁剪前,每个动作的内部动作平均替代值等于其TD信用。FACTOR在ALFWorld、WebShop和ScienceWorld上始终优于有竞争力的基线,每个环境种子的比较都支持FACTOR,且在最长视界环境上获得的增益最大。相同的超参数无需重新调整即可迁移到更大的骨干网络和不同的模型家族。 ablation研究表明,TD动作信用是改进的主要驱动因素,事后令牌分配则带来互补增益。

英文摘要:

Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distributing each action's credit across its tokens. In this paper, we introduce FACTOR, which separates these decisions. FACTOR uses checkpoint-calibrated TD residuals to assign per-action credits that telescope to the trajectory advantage, and feedback-conditioned teacher-student likelihood gaps to allocate each credit across the realized action tokens. Per-action normalization preserves the action-average coefficient and prevents token-level sign flips. We pair this construction with an action-mean reduction, removing the implicit dependence of an action's scalar surrogate weight on its token length. At the behavior policy and before clipping, each action's inner action-mean surrogate equals its TD credit. FACTOR consistently improves over competitive baselines across ALFWorld, WebShop, and ScienceWorld, with every environment-seed comparison favoring FACTOR and the largest gains emerging on the longest-horizon environment. The same hyperparameters transfer without retuning to a larger backbone and to a different model family. Ablations identify TD action credit as the dominant driver of the improvement, with hindsight token allocation contributing complementary gains.

↑