arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STAMP:深度搜索智能体的溯源引导信用分配

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

Ke Xu, Han Xu, Xinran Chen, Yuqian Wang, Zhixuan Li, Xiaojian Liu, Changwo Wu, Jianqiang Xia, Yuchen Li

arXiv 2607.11172首次发表:更新:

发表机构

Baidu Inc.; Peking University(百度公司; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究深度搜索智能体强化学习中奖励-信用不匹配问题,提出STAMP方法,通过基于参考的验证器和首次曝光归因确定信用,经符号保留优势调制注入信用,实验表明该方法能提升GRPO基线分数。

AI 中文摘要

深度搜索智能体的强化学习主要集中在轨迹级评分,如结果正确性、引用感知奖励和证据覆盖。然而,暴露支持文档的动作没有得到针对性的信用,即奖励-信用不匹配。我们提出了STAMP,其中基于参考的验证器判断每个引用文档是否支持训练时证据图中的实体或关系,首次曝光归因将每个支持的引用追溯到首次出现它的动作。通过符号保留优势调制注入步骤信用,在不改变轨迹级奖励或每组内轨迹相对排名的情况下重新分配优势。在BrowseComp、BrowseComp-ZH和xbench-DS上,STAMP在匹配的SFT初始化、训练数据和搜索工具下,将GRPO基线提高了+2.0/+5.5/+3.0分,并与仅结果和引用评分基础奖励相结合。组件消融证实了基于溯源的信用信号和符号保留优势调制各自对收益有贡献。

英文摘要

Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the reward-credit mismatch. We propose STAMP, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and first-exposure attribution traces each supported citation back to the action that first surfaced it. This step credit is injected through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group. On BrowseComp, BrowseComp-ZH, and xbench-DS, STAMP improves the GRPO baseline by +2.0/+5.5/+3.0 points under matched SFT initialization, training data, and search tools, and composes with both outcome-only and citation-rubric base rewards. Component ablations confirm that the provenance-based credit signal and the sign-preserving advantage modulation each contribute to the gains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑