PRO-Step:用于检索增强生成的步骤级过程奖励优化
PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation
浏览论文内容
中文总结 AI 辅助
针对RAG多跳推理的错误传播问题,提出PRO-STEP方法,通过训练生成式PRM、PRM引导的价值树搜索及步骤级DPO优化策略,在五个QA基准上取得最优平均EM和F1值。
中文摘要 AI 辅助
检索增强生成(Retrieval-Augmented Generation,RAG)通过将模型响应建立在外部知识基础上,增强了大型语言模型(Large Language Models,LLMs),但多跳推理仍易受错误传播影响,早期检索失败会混淆后续步骤。标准的基于结果的优化仅奖励最终答案,导致中间检索和推理错误未被检测到。现有基于过程的方法虽引入了步骤级信号,但仍将每个步骤与最终答案进行评分,奖励那些因有缺陷的检索偶然产生正确答案的虚假成功。RAG中的步骤级监督需要在每个步骤评估逻辑有效性和证据依据两个维度。我们提出PRO-STEP:训练一个生成式过程奖励模型(Process Reward Model,PRM),该模型可评估上述两个维度;采用PRM引导的价值树搜索构建偏好对,对比有效步骤与有缺陷步骤;通过步骤级直接偏好优化(Direct Preference Optimization,DPO)优化策略。在单跳和多跳问答数据集上的实验表明,PRO-STEP在五个基准测试中取得了最佳的精确匹配(Exact Match,EM)和F1值的平均值。代码、模型和训练数据可在该httpsURL公开获取。
英文摘要
Retrieval-Augmented Generation enhances Large Language Models by grounding responses in external knowledge, but multi-hop reasoning remains vulnerable to error propagation, where early retrieval failures confound subsequent steps. Standard outcome-based optimization only rewards the final answer, leaving intermediate retrieval and reasoning errors undetected. While existing process-based methods introduce step-level signals, they still score each step against the final answer, rewarding spurious successes where flawed retrieval coincidentally produces the correct answer. Step-level supervision in RAG requires evaluating both logical validity and evidential grounding at each step. We introduce PRO-STEP: we train a generative PRM that evaluates both dimensions, employ PRM-guided value tree search to construct preference pairs contrasting valid steps against flawed ones, and optimize the policy via step-level Direct Preference Optimization. Experiments on single and multi-hop QA datasets demonstrate that PRO-STEP achieves the best average EM and F1 across five benchmarks. Code, models, and training data are publicly available at https://github.com/keemminnke/PRO-Step.
发表机构
- Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。