arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SLCA-GRPO:解决工具调用强化学习中的跨段信用分配错误

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

Yan Zhan, Shaobo Liu, Qiunan Liu, Yuanjun Shi, Siqi Xu, WeiYi Hou, Xiang Xu, Zekang Li, Weizhou Pan, Jiahong Yan

arXiv 2609.29050首次发表:更新:

发表机构

Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对工具调用强化学习中跨段信用分配错误,提出SLCA-GRPO框架,通过段锁定信用分配和分层奖励解耦优势,在7B模型上显著提升性能。

AI 中文摘要

工具调用智能体产生异构输出,将结构化工具调用与面向用户的自然语言摘要交错在一起。这种输出异构性在标准的同策略强化学习(RL)中呈现出一个结构性失败模式:像GRPO这样的算法不加区分地将同质的轨迹级标量优势广播给所有令牌。因此,来自摘要生成的梯度噪声泄漏到工具决策令牌中,导致跨段信用分配错误和脆弱的优化。在这项工作中,我们提出了SLCA-GRPO,一个包含段锁定信用分配(SLCA)的框架。为了在没有昂贵真实API的情况下实现可扩展的探索和稳定的训练,我们首先构建了模式引导的LLM模拟器(SGLS)作为基础训练基础设施。在此基础上,SLCA在单组rollout中在结构段级别解耦优势估计,无需从中间状态进行额外的rollout。在分层奖励(HierR)的支持下,SLCA将执行优势路由到工具令牌,将偏好优势路由到摘要令牌,消除了每次策略更新中的优势污染(跨段信用分配错误的主要渠道)。在7B骨干模型上,SLCA-GRPO加速了收敛,并在相同训练预算下,在域内评估上比标准GRPO、ToolPO和RLTR高出+2.53个百分点,在伯克利函数调用排行榜(BFCL)上高出+1.36个百分点,在τ²-Bench上高出+9.15个百分点,实现了更高的准确性,同时减少了工具冗余和成本。

英文摘要

Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In this work, we propose SLCA-GRPO, a framework incorporating Segment-Locked Credit Assignment (SLCA). To enable scalable exploration without costly real APIs and stable training, we first construct the Schema-Guided LLM Simulator (SGLS) as foundational training infrastructure. Building on this, SLCA decouples advantage estimation at the structural segment level within a single group of rollouts, without requiring additional rollouts from intermediate states. Supported by Hierarchical Rewards (HierR), SLCA routes execution advantages to tool tokens and preference advantages to summary tokens, eliminating advantage contamination (the dominant cross-segment credit misattribution channel) within each policy update. On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in-domain evaluation, +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL), and +9.15 pp on $τ^2$-Bench under the same training budgets, achieving higher accuracy with reduced tool redundancy and costs.

Comments41 pages, 13 figures. Code: https://github.com/SLCA-GRPO/SLCA-GRPO ; Dataset: https://huggingface.co/datasets/YanZhanPKU/SLCA-GRPO-Datasets

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑