arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14747cs.LG

WANDR:宽深研究基准

WANDR: A Benchmark for Wide and Deep Research

Vitaliy Polshkov, Marcin Pitera, Jeremy Yang, Kirill Priemko, Maksim Gaiduk, Aleksandr Nikolenko, Denis Bykov, Clare Southern, Denis Yarats, Jerry Ma

AI总结:

该研究提出WANDR基准,含500个现实数据收集任务,用于评估研究智能体,发现现有系统性能远未饱和,性能受目标量和层级深度影响,存在发现不完整等瓶颈。

AI中文摘要:

WANDR(Wide ANd Deep Research,宽深研究)是一个包含500个现实且具有挑战性的数据收集任务的研究智能体基准。每个任务要求系统发现满足指定标准的大量实体(广度要求),通过多次协调的网络搜索调查每个实体(深度要求),并返回带有支持来源和摘录的可独立验证记录。任务以资格键层级结构表示,该结构指定实体、关系、证据以及每个层级所需的数量;具有n家公司、每家公司m名员工、每名员工k个来源的层级结构需要n×m×k条记录。该结构支持市场映射、尽职调查、文献综述、产品比较和人才挖掘等多种工作流程,目标记录数量从数十到数千不等。WANDR用特定任务的评判者取代静态标准答案集,这些评判者会重新获取引用的页面并根据证据验证每条记录,从而能够评估当前及不断变化的事实。记录的判定结果被汇总为软、硬精确率、召回率和F1分数,以区分事实质量、覆盖范围和层级完整性。这些任务来自去标识化的产品使用日志,通过半自动化流程生成,该流程包含自动检查、实证审计以及必要时的人工审核。我们评估了6个生产级研究系统,发现该基准远未达到饱和:在高投入下,最强的系统仅达到0.363的软F1分数和0.133的硬F1分数。性能会随目标数量和层级深度的增加而下降,未完成的发现、缺失的丰富度以及不完整的证据构建仍是主要瓶颈。该基准和评估工具可在指定的https URL获取。

英文摘要:

WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set of entities that satisfy specified criteria (breadth), investigate each entity through multiple coordinated web searches (depth), and return independently verifiable records with supporting sources and excerpts. Tasks are represented as qualification key hierarchies that specify the entities, relationships, evidence, and required count at each level; a hierarchy with n companies, m employees per company, and k sources per employee requires n x m x k records. This structure supports diverse workflows such as market mapping, due diligence, literature review, product comparison, and talent sourcing, with targets ranging from dozens to thousands of records. WANDR replaces static gold answer sets with task-specific judges that refetch cited pages and verify each record against its evidence, allowing evaluation of current and changing facts. Record verdicts are aggregated into soft and hard precision, recall, and F1 scores that distinguish factual quality, coverage, and hierarchical completeness. The tasks are derived from de-identified product-usage logs and produced through a semi-automated pipeline with automated checks, empirical audits, and human review where needed. We evaluate six production research systems and find that the benchmark is far from saturated: at high effort, the strongest system reaches only 0.363 soft F1 and 0.133 hard F1. Performance degrades as target volume and hierarchy depth increase, with incomplete discovery, missing enrichment, and incomplete evidence construction remaining major bottlenecks. The benchmark and evaluation harness are available at https://github.com/perplexityai/wandr.

↑