发表机构
Perplexity AI(Perplexity AI公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Q2D-Web是一个大规模智能体检索基准,含1.9亿文档和7万查询,提供三组相关性判定,评估13种检索器,并验证子语料库采样可近似全语料库评估。
AI 中文摘要
在大规模生产级RAG系统中评估第一阶段检索器,需要一个基准,该基准将大规模语料库与大量基于真实用户查询及其对话线程的智能体重构搜索查询配对,并为每个查询标注多个相关文档。现有的公开基准均未评估此设置:大规模集合通常仅提供少量评估查询,而包含大量查询的基准通常仅包含数百万篇文档。此外,大多数基准评估的是人类撰写的查询,而智能体RAG流水线中的第一阶段检索器服务于机器撰写的重构查询,其分布与人类搜索行为不同。为克服这些评估空白,我们提出了Q2D-Web(Query2Doc-Web),一个大规模智能体检索基准,包含1.9亿文档的网络语料库和70,000个十种语言的智能体搜索查询,这些查询从生产系统中的真实用户查询重构而来。Q2D-Web提供了三组固定的相关性判定:智能体引用、生产排名以及组合集,该组合集合并了两种信号,并添加了基于LLM的未标注池化文档判定以减少假阴性。我们评估了13种检索器,包括词法、稠密和后期交互模型,发现它们的相对排序在很大程度上对判定集的选择不敏感,但在主题领域、查询语言和查询类型上存在显著差异。为支持快速评估,我们还研究了子语料库采样作为全语料库评估的近似方法。保留三分之一的语料库,通过池化检索器运行结果的倒数排名融合进行选择,在组合判定下保持了全语料库的模型排名,同时仅将绝对Recall@1000提高了3到7个百分点。公开排行榜可通过以下网址访问:此https URL
英文摘要
Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant documents per query. No existing public benchmark evaluates this setting: large-scale collections typically provide only a small number of evaluation queries, whereas benchmarks with many queries generally contain only millions of documents. Moreover, most benchmarks assess human-written queries, while the first-stage retrievers in agentic RAG pipelines serve machine-written reformulations whose distribution differs from human search behavior. To overcome these evaluation gaps, we introduce Q2D-Web (Query2Doc-Web), a large-scale agentic retrieval benchmark consisting of a 190M-document web corpus and 70k agentic search queries in ten languages, reformulated from real-world user queries in production systems. Q2D-Web provides three sets of fixed relevance judgments: agent citations, production rankings, and a combined set that unions both signals and adds LLM-based judgments of unlabeled pooled documents to reduce false negatives. We benchmark 13 retrievers including lexical, dense, and late-interaction models and find that their relative ordering is largely insensitive to the choice of judgment set, while diverging substantially across topical domains, query languages, and query types. To enable fast evaluation, we also study subcorpus sampling as an approximation to full-corpus evaluations. Retaining a third of the corpus, selected by reciprocal rank fusion over pooled retriever runs, preserves the full-corpus model ranking under the combined judgments while raising absolute Recall@1000 only by 4 to 7 points. The public leaderboard is accessible under: https://huggingface.co/spaces/perplexity-ai/q2d-web-leaderboard