将BrowseComp-Plus映射到ClimbMix:构建更符合智能体搜索需求的真实语料库
Projecting BrowseComp-Plus onto ClimbMix: Toward More Realistic Corpora for Agentic Search
AI总结:
该研究将BrowseComp-Plus的查询映射到NVIDIA的ClimbMix语料库,构建了与基准无关的映射流水线,生成57个有效查询,使智能体搜索难度向检索环节转移,发布了相关资源。
AI中文摘要:
BrowseComp-Plus基准通过将不透明的网络搜索替换为固定语料库,拆分了智能体搜索的评估,使智能体的角色与检索器的角色得以分离。然而,该语料库仅包含约10万份文档,由基准自身查询的支持文档加上挖掘的困难负例组装而成,因此证据和干扰项均是针对每个查询单独选择的。我们引入了BrowseComp-Plus_CM,它保留了BrowseComp-Plus的查询,但将其证据迁移至ClimbMix——这是NVIDIA发布的用于预训练语言模型的4000亿token、5.53亿份文档的网络文本混合语料库,且构建过程未参考任何基准。我们的主要贡献是实现这一迁移的映射流水线:它将每个查询分解为原子推理步骤,并将每个步骤锚定到新语料库中,仅当自动验证、独立智能体和人工审核均确认每个步骤都有支撑时,才保留该查询。该流水线与数据集无关,适用于任何查询可分解为可验证事实的基准。将其应用于830个BrowseComp-Plus测试查询时,该流水线生成了57个完全锚定且带有查询级相关性判断的查询。迁移将难度转移到了检索环节:我们评估的最强智能体的答案准确率下降了5个百分点,证据召回率从84.3%降至21.4%,同时发出的搜索调用增加了63%。作为一系列映射工作的首个成果,我们在该httpsURL发布了该流水线、基准及相关分析。
英文摘要:
The BrowseComp-Plus benchmark disentangled the evaluation of agentic search by replacing opaque web search with a fixed corpus, so that an agent's role can be separated from the retriever's. That corpus, however, holds only about 100K documents and was assembled from the supporting documents of the benchmark's own queries plus mined hard negatives, so the evidence and the distractors were both selected per query. We introduce $\text{BrowseComp-Plus}_{\text{CM}}$, which keeps the BrowseComp-Plus questions but relocates their evidence to ClimbMix, a 400B-token, 553M-document mixture of web text released by NVIDIA for pre-training language models and built without reference to any benchmark. Our main contribution is the projection pipeline that makes this possible: it decomposes each question into atomic reasoning hops and grounds every hop in the new corpus, retaining a question only when automatic verification, an independent agent, and human review all confirm that every hop is supported. The pipeline is dataset-agnostic and applies to any benchmark whose questions decompose into verifiable facts. Applied to the 830 BrowseComp-Plus test questions, our pipeline yields 57 fully grounded questions with question-level relevance judgments. Projection shifts the difficulty onto retrieval, as the strongest agent we evaluate loses five points of answer accuracy but sees its evidence recall fall from 84.3% to 21.4% while issuing 63% more search calls. As the first of a series of projections, we release the pipeline, the benchmark, and our analyses at https://github.com/castorini/cmass.