arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02163cs.AI

从简单问答到深度研究:通过迭代任务演化构建的可验证基准

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

Can Wang, Haoran Chen, Haowen Gao, Hao Ding, Zhaoyang Liu, Zhiying Tu

AI总结:

该研究针对现有深度研究基准的验证缺陷,通过迭代流水线自动构建含500个任务的可验证基准,可区分模型与查询类型,支持稳定评估,相关数据等公开可用。

AI中文摘要:

深度研究基准需要专家级任务和基于特定任务知识的可靠评估。现有基准严重依赖专家编写或预先存在的人工编写材料,而全自动构建难以确保一致且可追踪的验证。为解决这一差距,我们引入了一个包含500个深度研究任务的可验证基准,涵盖31个主题和10个主要类别,设计了三种查询形式以探测深度研究所需的互补能力。该基准使用迭代的Explorer-Formalizer-Challenger流水线自动构建,该流水线逐步将简单问题转化为深度研究任务。每个任务表示为原子步骤和相关检查点的有向无环图(DAG),使查询、DAG和评分标准能以可控方式共同演化。实验表明,该基准可清晰区分模型和查询类型,而其基于事实的逐点评分标准支持细粒度、符合人类偏好且稳定的评估。我们的数据、实现和结果均公开可用。

英文摘要:

Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existing human-authored materials, while fully automatic construction struggles to ensure consistent and traceable verification. To address this gap, we introduce a verifiable benchmark of 500 deep research tasks spanning 31 topics and 10 major categories, with three query forms designed to probe complementary capabilities required for deep research. The benchmark is constructed automatically using an iterative Explorer-Formalizer-Challenger pipeline that progressively transforms simple questions into deep research tasks. Each task is represented as a directed acyclic graph (DAG) of atomic steps and associated checkpoints, enabling the query, DAG, and rubrics to evolve together in a controlled manner. Experiments demonstrate that the benchmark clearly discriminates among models and query types, while its fact-grounded pointwise rubrics enable fine-grained, human-aligned, and stable evaluation. Our data, implementation, and results are publicly available.

补充信息

↑