arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19827cs.CL

F$^{2}$DR:面向DeepSearch工作流的细粒度全流水线奖励框架

F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

Bojian Xiong, Wentao Ding, Yujing Lu, Shaowei Zhang, Ling Shi, Jing Liao, Yan Wang, Yueyang Zhang, Long Xia, Zhiyuan Sun, Daiting Shi, Jingzhou He, Yuqi Ren, Deyi Xiong

首次发表
浏览论文内容

中文总结 AI 辅助

针对DeepSearch工作流缺乏全流水线奖励评估的问题,提出F2DR框架,从内容、轨迹和答案三维度评估,并构建DeepSearch RM-Bench基准,实验证明其评估一致性与区分能力显著优于基线。

中文摘要 AI 辅助

随着大型语言模型(LLMs)在工业界的广泛部署,DeepSearch已成为解决复杂用户查询的主流范式。它通常通过一个由规划与反思、信息检索和答案生成组成的迭代闭环工作流来运行。然而,现有的奖励模型(RMs)和评估基准主要针对静态单轮任务设计,无法捕捉DeepSearch工作流的全流水线复杂性。为解决这一局限,我们提出了F2DR,一个细粒度的全流水线DeepSearch奖励框架。F2DR从三个维度评估DeepSearch工作流:内容(Content)、轨迹(Trajectory)和答案(Answer),从而实现全面的过程级评估。我们进一步构建了DeepSearch RM-Bench,一个专门用于在DeepSearch场景中评估奖励模型的基准。大量实验表明,F2DR相比基于自我评估的基线取得了显著更高的评估一致性,而DeepSearch RM-Bench在现有开源奖励模型上展现出强大的区分能力。我们将很快公开发布完整的DeepSearch RM-Bench数据集。

英文摘要

With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, information retrieval, and answer generation. However, existing reward models (RMs) and evaluation benchmarks are primarily designed for static single-turn tasks, failing to capture the full-pipeline complexity of DeepSearch workflows. To address this limitation, we propose F2DR, a fine-grained full-pipeline DeepSearch reward framework. F2DR evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer, enabling comprehensive process-level assessment. We further construct DeepSearch RM-Bench, a dedicated benchmark for evaluating RMs in DeepSearch scenarios. Extensive experiments demonstrate that F2DR achieves significantly higher evaluation consistency than self-evaluation-based baselines, while DeepSearch RM-Bench exhibits strong discriminative capability across existing open-source RMs. We will publicly release the complete DeepSearch RM-Bench dataset soon.

发表机构

  • Tianjin University(天津大学)
  • Baidu Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

↑