arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mr.LHDR:多模态真实世界长时程深度研究智能体基准

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Minghao Guo, Meng Cao, Sui Zhao, Siyu Ning, Xin Wang, Haoze Zhao, Jiaxuan Yang, Haihong Hao, Mingfei Han, Shunlin Rong, Haijun Wu, Xiaodan Liang, Xiaojun Chang

arXiv 2609.11318首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; University of Science and Technology of China; Zhejiang University; Tencent(穆罕默德·本·扎耶德人工智能大学; 中国科学技术大学; 浙江大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Mr.LHDR基准,用隐藏图构建长依赖链问题,评估多模态深度研究智能体,发现最强系统OA仅43.1%,证据整合是瓶颈。

AI 中文摘要

深度研究智能体在网络搜索、工具使用、多模态证据分析和信息综合方面的能力日益增强。然而,现有基准主要评估中等时程的探索,很少测试智能体能否维持长时程、依赖关系繁重的研究过程。我们提出了Mr.LHDR(多模态真实世界长时程深度研究),这是一个用于评估在八个类别中,对长时程、不可约简的相互依赖证据链进行真实世界深度研究的基准。每个问题都基于一个隐藏的节点-关系图构建,在得出简短、唯一且可验证的答案之前,平均需要12.1个必要的中期结论,平均依赖深度为10.4。问题包含多模态证据,包括图像、地图、PDF、标志、图表、表格和视频帧,并且至少有一个非文本元素会改变推理状态。Mr.LHDR在标注的依赖关系下,同时评估最终答案和中期结论的正确性。我们使用总体准确率(OA)、严格准确率(SA)、检查表得分(CS)和依赖感知检查表得分(DACS)来评估通用模型、深度研究系统和智能体框架。结果表明,即使最强的系统也仅达到43.1%的OA和34.3%的SA,这表明最终答案的准确率大大高估了完整的研究成功。移除图像会使DACS降低12.6分,证明了多模态证据的重要性,而随着推理链变长,SA持续下降。这些发现揭示了持续的、依赖一致的证据整合(而非孤立的事实检索)是当前深度研究智能体的关键瓶颈。

英文摘要

Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes. We introduce Mr. LHDR (Multimodal real-world Long-Horizon Deep Research), a benchmark for evaluating real-world deep research over long, irreducible chains of interdependent evidence across eight categories. Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer. Questions incorporate multimodal evidence, including images, maps, PDFs, logos, charts, tables, and video frames, with at least one non-text element that changes the reasoning state. Mr. LHDR evaluates both final answers and the correctness of intermediate conclusions under annotated dependencies. We evaluate general models, deep research systems, and agent frameworks using Overall Accuracy (OA), Strict Accuracy (SA), Checklist Score (CS), and Dependency-Aware Checklist Score (DACS). Results show that even the strongest system achieves only 43.1% OA and 34.3% SA, indicating that final-answer accuracy substantially overestimates complete research success. Removing images reduces DACS by 12.6 points, demonstrating the importance of multimodal evidence, while SA consistently declines as reasoning chains become longer. These findings reveal sustained, dependency-consistent evidence integration, rather than isolated fact retrieval, as a key bottleneck for current deep research agents.

CommentsCode and data are available at https://github.com/minghaoguo20/Mr-LHDR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑