面向深度研究智能体的、基于搜索 rubrics 的训练文档重排序器
Training Documents Reranker with Search Rubrics for Deep Research Agent
浏览论文内容
中文总结 AI 辅助
针对现有检索器无法满足深度研究智能体复杂信息需求的问题,提出搜索 rubrics 并训练文档重排序器 RubricRanker,在多基准上取得优于基线的性能且泛化性良好。
中文摘要 AI 辅助
检索系统通过提供相关文档帮助深度研究智能体生成高质量答案,但现有检索器通常通过相关性匹配选择文档,而单独匹配度高的前 k 个文档可能无法形成满足智能体查询复杂信息需求(如多样、简洁且权威的文档)的集合。本文提出搜索导向 rubrics,为每个智能体查询明确界定高质量文档集合应满足的要求,这些搜索 rubrics 被组织为分层结构并通过强大的大语言模型(LLM)合成。基于这些搜索 rubrics,我们进一步训练文档重排序器 RubricRanker,从检索到的文档中选择高质量子集。我们设计了包含 rubrics 引导的监督微调与基于 rubrics 的强化学习的两阶段训练框架。大量实验表明,RubricRanker 在四个深度研究基准上比最强基线高出 2.6 个百分点,且能良好泛化到五个检索增强生成(RAG)基准。
英文摘要
Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically select documents through relevance matching, while individually well-matched top-$k$ documents may not form a \textit{set} that satisfies the complex information needs of an agent query (\eg, diverse, concise and authoritative documents). In this paper, we propose search-oriented rubrics that \textit{explicitly} define the requirements that high-quality document sets should satisfy for each agent query. Our search rubrics are organized into a hierarchical structure and synthesized using a powerful LLM. Based on these search rubrics, we further train a document reranker \textbf{RubricRanker} to select a high-quality subset from retrieved documents. We design a two-stage training framework that consists of rubrics-guided supervised fine-tuning and rubric-based reinforcement learning. Extensive experiments demonstrate that RubricRanker outperforms the strongest baseline by 2.6 points on four deep research benchmarks and generalizes well to five RAG benchmarks.
发表机构
- Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。