发表机构
State Key Laboratory for Novel Software Technology, Nanjing University; Ant Group; National Institute of Healthcare Data Science, Nanjing University(南京大学计算机软件新技术国家重点实验室; 蚂蚁集团; 南京大学国家健康医疗数据科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对搜索智能体强化学习的信用分配问题,提出基于事实效用估计的密集过程监督方法,在7个QA基准上优于现有基线。
AI 中文摘要
搜索智能体的强化学习(RL)通常依赖于结果奖励,但由于中间步骤的价值不明确,往往无法实现有效的信用分配,难以将其贡献与最终结果分离开来。本文提出一种基于事实效用估计的密集过程监督方法,将推理过程建模为离散证据事实的累积。首先从原始观测中提取结构化事实并组织为显式事实库;为支持信用分配,对语义等价事实进行聚类,通过群体rollout的贝叶斯估计推断每个事实簇的后验效用;最后将估计的事实效用转换为密集的步骤级奖励以指导RL训练。在7个单跳和多跳QA基准上的实验表明,该方法始终优于现有基线, ablation研究验证了其与仅使用结果奖励训练相比,在多跳QA上的显著相对改进。
英文摘要
Reinforcement learning (RL) for search agents typically relies on outcome rewards. However, it often fails to achieve effective credit assignment, due to the unclear value of intermediate steps. It is hard to separate their contributions from the final result. In this paper, we propose a dense process supervision method based on fact utility estimation, which models the reasoning process as the accumulation of discrete evidence facts. We first extract structured facts from raw observations and organize them into an explicit fact store. To support credit assignment, we then cluster semantically equivalent facts and infer the posterior utility of each fact cluster using Bayesian estimation over group rollouts. Finally, we convert the estimated fact utilities into dense step-level rewards to guide RL training. Experiments on seven single-hop and multi-hop QA benchmarks show that our method consistently outperforms existing baselines. Ablation studies validate clear relative improvements on multi-hop QA compared to outcome reward-only training.
CommentsAccepted in the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)