arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HiEviDR-Bench:深度研究中分层证据聚合的基准测试

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research

Yubo Sun, Chunyi Peng, Yukun Yan, Zhenghao Liu, Sen Mei, Bangrui Xu, Xuanhe Zhou, Chi Chen, Maosong Sun

arXiv 2607.25151首次发表:更新:

AI 中文总结

针对深度研究中证据聚合评估不足问题,HiEviDR-Bench构建基准测试,涵盖多领域多模态设置,用证据图表示实例,开发评估框架及机制。含2000个问题,实验显示多模型虽报告质量好,但在引用准确性等方面欠佳,瓶颈在证据识别和中间主张构建。

AI 中文摘要

深度研究要求模型从大规模异构源中检索、连接和综合证据,以回答复杂查询并生成分析报告。现有基准主要评估最终结果,对证据选择、链接和聚合过程缺乏洞察。为此引入HiEviDR-Bench,涵盖多种设置,用证据图表示实例。开发了可追溯性评估框架及渐进门控机制。该基准含2000个人工验证问题,实验表明多模型在报告质量上佳,但在引用准确性等方面表现不佳,瓶颈在于证据识别和中间主张构建。

英文摘要

Deep research requires models to retrieve, connect, and synthesize evidence from large-scale heterogeneous sources to answer complex queries and produce analytical reports. Existing benchmarks mainly evaluate final outcomes, such as answer correctness, report quality, or citation alignment, while providing limited visibility into whether evidence is correctly selected, linked, and aggregated into supported claims and conclusions. To address this gap, we introduce HiEviDR-Bench, a benchmark for evaluating Hierarchical Evidence Aggregation in Deep Research. HiEviDR-Bench covers open-domain and academic-domain settings under both text-only and multimodal conditions, and represents each instance with an explicit evidence graph that captures evidence selection, cross-source linking, and aggregation from evidence to intermediate claims and final conclusions. Based on this formulation, we develop a traceability-oriented evaluation framework with five dimensions: report quality, evidence traceability, citation accuracy, claim verification, and answer correctness, together with a progressive gating mechanism for fine-grained error localization. HiEviDR-Bench contains 2,000 human-validated questions with evidence graphs across multiple difficulty levels. Experiments on 16 representative multimodal large language models show that, although many systems achieve strong report quality, their performance drops markedly on citation accuracy, claim construction, and answer correctness. Further analysis shows that the main bottlenecks lie in evidence identification and intermediate claim construction, revealing that strong surface-level report quality does not necessarily imply grounded multi-stage reasoning on our benchmark.

CommentsCode and data are available at https://ai9stars.github.io/HiEviDR-Bench.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑