arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越以相关性为中心的检索:面向评分标准的文档集选择与排序

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

Kailin Jiang, Lei Liu, Jian Xi, Hui Xu, Junlin Liu, Baochen Fu, Bin Li, Vichwang, Yu Lu, Haibo Shi

arXiv 2607.19747首次发表:更新:

发表机构

University of Science and Technology of China; Yuanbao Team, Tencent; University of Chinese Academy of Sciences; Shandong University(中国科学技术大学; 腾讯元宝团队; 中国科学院大学; 山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型和人工智能代理对搜索结果质量需求,提出评估-诊断-优化框架。设计SetwiseEvalKit评估基准,评估多种重排器。在此基础上提出Rubric4Setwise方法,无需训练,能以更少文档和搜索轮次实现最佳下游生成性能,验证了闭环有效性。

AI 中文摘要

随着大语言模型和人工智能代理成为搜索结果的主要消费者,文档集质量决定了下游生成的上限。然而,现有评估系统仍局限于独立给文档打分并通过nDCG聚合,忽略文档间交互,无法回答为何一个文档集优于另一个。为解决这些问题,我们提出一个完整的评估-诊断-优化框架。设计了SetwiseEvalKit,一个涵盖短文本和长文本场景的三级九维文档集评估基准,包含约28K高质量评估标准。系统评估了12种重排器,最佳方法覆盖率不超45%,跨文档协调维度普遍较弱。在此基础上,提出Rubric4Setwise,一种无需训练的方法,将基于评分标准的评估标准转换为文档集选择信号,用更少文档和搜索轮次实现最佳下游生成性能,是唯一在两种场景下均保持最优结果的方法,验证了从评估到优化闭环的有效性。

英文摘要

As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, conflict, complementarity) and unable to answer what makes one document set better than another. To address these issues, we propose a complete evaluate-diagnose-optimize framework. We design SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark covering both short-form and long-form scenarios, comprising approximately 28K high-quality evaluation rubrics. We systematically evaluate 12 rerankers: even the best method achieves no more than 45% coverage, cross-document coordination dimensions are universally weak, and no single method maintains top performance across both settings. Building on this, we propose Rubric4Setwise, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds. It is the only method that maintains state-of-the-art results across both scenarios, validating the effectiveness of closing the loop from evaluation to optimization.

CommentsProject Page: https://rubric4setwise.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑