F$^{2}$DR:面向DeepSearch工作流的细粒度全流水线奖励框架
F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows
浏览论文内容
中文总结 AI 辅助
针对DeepSearch工作流缺乏全流水线奖励评估的问题,提出F2DR框架,从内容、轨迹和答案三维度评估,并构建DeepSearch RM-Bench基准,实验证明其评估一致性与区分能力显著优于基线。
中文摘要 AI 辅助
随着大型语言模型(LLMs)在工业界的广泛部署,DeepSearch已成为解决复杂用户查询的主流范式。它通常通过一个由规划与反思、信息检索和答案生成组成的迭代闭环工作流来运行。然而,现有的奖励模型(RMs)和评估基准主要针对静态单轮任务设计,无法捕捉DeepSearch工作流的全流水线复杂性。为解决这一局限,我们提出了F2DR,一个细粒度的全流水线DeepSearch奖励框架。F2DR从三个维度评估DeepSearch工作流:内容(Content)、轨迹(Trajectory)和答案(Answer),从而实现全面的过程级评估。我们进一步构建了DeepSearch RM-Bench,一个专门用于在DeepSearch场景中评估奖励模型的基准。大量实验表明,F2DR相比基于自我评估的基线取得了显著更高的评估一致性,而DeepSearch RM-Bench在现有开源奖励模型上展现出强大的区分能力。我们将很快公开发布完整的DeepSearch RM-Bench数据集。
英文摘要
With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, information retrieval, and answer generation. However, existing reward models (RMs) and evaluation benchmarks are primarily designed for static single-turn tasks, failing to capture the full-pipeline complexity of DeepSearch workflows. To address this limitation, we propose F2DR, a fine-grained full-pipeline DeepSearch reward framework. F2DR evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer, enabling comprehensive process-level assessment. We further construct DeepSearch RM-Bench, a dedicated benchmark for evaluating RMs in DeepSearch scenarios. Extensive experiments demonstrate that F2DR achieves significantly higher evaluation consistency than self-evaluation-based baselines, while DeepSearch RM-Bench exhibits strong discriminative capability across existing open-source RMs. We will publicly release the complete DeepSearch RM-Bench dataset soon.
发表机构
- Tianjin University(天津大学)
- Baidu Inc.(百度公司)
机构由 AI 辅助整理,请以论文原文为准。