arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13501cs.AI

LAPO:多轮搜索推理中自生成过程奖励的留一回合归因

LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning

Qiang Zhu, Jiajun Wu, Longyi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究多轮搜索推理中强化学习依赖终端结果奖励的问题,提出基于反向留一回合归因的LAPO方法,无需额外模型,在七个问答数据集上取得较好成绩,证明策略追溯归因能为多轮搜索智能体提供有效过程监督。

中文摘要 AI 辅助

多轮搜索推理的强化学习通常依赖终端结果奖励,无法区分有用、冗余和有害的中间交互。我们提出LAPO,一种基于反向留一回合归因的自生成过程监督方法。对于每个搜索回合,LAPO用固定占位符替换该回合及其检索观察,并测量当前策略对正确答案的平均对数似然的变化。答案似然增益估计该回合的贡献,同时保留所有下游交互。LAPO还应用符号一致性门控,仅保留方向与其原始归因分数一致的归一化过程优势。该方法无需额外奖励模型等。在七个带局部检索的知识密集型问答数据集上,LAPO平均精确匹配分数达0.326,优于最强步骤奖励基线IGPO 0.053。消融实验表明反向归因和符号一致性门控的互补益处,证明基于策略的追溯归因可为多轮搜索智能体提供有效过程监督。

英文摘要

Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactions. We propose LOTAPO , a self-generated process-supervision method based on backward leave-one-turn attribution. For each search turn, LOTAPO replaces the turn and its retrieval observation with a fixed [DELETE] placeholder and measures the resulting change in the current policy's mean log-likelihood of the gold answer. This Answer-Likelihood Gain estimates the turn's contribution while preserving all downstream interactions, allowing early evidence to be evaluated in the complete reasoning context. LOTAPO further applies sign-consistency gating, retaining only normalized process advantages whose directions agree with their raw attribution scores. The method requires no additional reward model, teacher, verifier, or LLM-as-a-Judge. Across seven knowledge-intensive question-answering datasets with local retrieval, LOTAPO achieves an average exact-match score of 0.326, outperforming the strongest step-reward baseline, IGPO, by 0.053. Ablations show complementary benefits from backward attribution and sign-consistency gating, demonstrating that policy-derived retrospective attribution can provide effective process supervision for multi-turn search agents.

发表机构

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑