arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03527cs.IRcs.AIcs.CL

面向深度研究智能体的、基于搜索 rubrics 的训练文档重排序器

Training Documents Reranker with Search Rubrics for Deep Research Agent

Wenhan Liu, Yu Lu, Qiaolin Xia, Hui Xu, Tong Zhao, Jian Xi, Yutao Zhu, Haijin Liang, Haibo Shi, Hao Wang, Zhicheng Dou

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有检索器无法满足深度研究智能体复杂信息需求的问题,提出搜索 rubrics 并训练文档重排序器 RubricRanker,在多基准上取得优于基线的性能且泛化性良好。

中文摘要 AI 辅助

检索系统通过提供相关文档帮助深度研究智能体生成高质量答案,但现有检索器通常通过相关性匹配选择文档,而单独匹配度高的前 k 个文档可能无法形成满足智能体查询复杂信息需求(如多样、简洁且权威的文档)的集合。本文提出搜索导向 rubrics,为每个智能体查询明确界定高质量文档集合应满足的要求,这些搜索 rubrics 被组织为分层结构并通过强大的大语言模型(LLM)合成。基于这些搜索 rubrics,我们进一步训练文档重排序器 RubricRanker,从检索到的文档中选择高质量子集。我们设计了包含 rubrics 引导的监督微调与基于 rubrics 的强化学习的两阶段训练框架。大量实验表明,RubricRanker 在四个深度研究基准上比最强基线高出 2.6 个百分点,且能良好泛化到五个检索增强生成(RAG)基准。

英文摘要

Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically select documents through relevance matching, while individually well-matched top-$k$ documents may not form a \textit{set} that satisfies the complex information needs of an agent query (\eg, diverse, concise and authoritative documents). In this paper, we propose search-oriented rubrics that \textit{explicitly} define the requirements that high-quality document sets should satisfy for each agent query. Our search rubrics are organized into a hierarchical structure and synthesized using a powerful LLM. Based on these search rubrics, we further train a document reranker \textbf{RubricRanker} to select a high-quality subset from retrieved documents. We design a two-stage training framework that consists of rubrics-guided supervised fine-tuning and rubric-based reinforcement learning. Extensive experiments demonstrate that RubricRanker outperforms the strongest baseline by 2.6 points on four deep research benchmarks and generalizes well to five RAG benchmarks.

发表机构

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑