混合深度研究在数据库查询与网络搜索中的基准测试
Benchmarking Hybrid Deep Research Across Database Querying and Web Search
- University of Houston(休斯顿大学)
- University of California, Los Angeles(加州大学洛杉矶分校)
- The University of Hong Kong(香港大学)
- Seoul National University(首尔大学)
- Snowflake AI Research(Snowflake AI研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有深度研究基准孤立评估网络搜索与数据库查询的不足,提出HybridDeepResearch基准,要求智能体结合两者完成380个任务,发现最先进模型在困难子集上Pass@8仅约50-54%,揭示跨模态约束保持仍是开放挑战。
AI中文摘要:
尽管自主智能体在“深度研究”方面已取得显著进展,能够通过迭代浏览开放网络来综合信息,但现实世界中的问题解决很少局限于单一环境。复杂的分析任务本质上要求智能体将来自模糊非结构化文本(如开放网络)和高度精确结构化数据(如关系数据库)的证据编织在一起。然而,现有基准测试孤立地评估这些模态,未能捕捉关键的“交接”——即在系统之间移动证据时保持约束的能力。我们引入了HybridDeepResearch,据我们所知,这是首个要求同时使用网络搜索和SQL来形成完整、可验证答案的深度研究基准。该基准包含380个基于LiveSQLBench-Base-Lite数据库和公共网络语料库的工具依赖任务,通过自动检查和人工审查验证,涵盖三种推理模式:SQL2S、S2SQL和Parallel。在多种智能体框架下对专有和开放权重模型的评估显示,即使像GLM-5.2、Claude-Sonnet-4.6和GPT-5这样的最先进模型,在困难子集上的Pass@8也仅达到约50-54%。值得注意的是,结果表明定向推理比并行交集困难得多,突显了在不丢失约束的情况下桥接结构化与非结构化信息空间,对智能体系统来说仍是一个重大的开放挑战。代码和数据集已在GitHub(此https URL)和Hugging Face(此https URL)上公开提供。
英文摘要:
While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave together evidence from both ambiguous unstructured text (e.g., the open web) and highly precise structured data (e.g., relational databases). However, existing benchmarks evaluate these modalities in isolation, failing to capture the critical "handoff" - the ability to preserve constraints when moving evidence between systems. We introduce HybridDeepResearch, to our knowledge the first deep-research benchmark that requires both web search and SQL to form a complete, verifiable answer. The benchmark contains 380 tool-dependent tasks grounded in LiveSQLBench-Base-Lite databases and public web corpora, validated through automated checks and human review, and covering three reasoning patterns: SQL2S, S2SQL, and Parallel. Evaluations across proprietary and open-weight models under various agentic scaffolds reveal that even state-of-the-art models like GLM-5.2, Claude-Sonnet-4.6 and GPT-5 achieve only about 50-54% Pass@8 on the hard subset. Notably, results show that directional reasoning is substantially more difficult than parallel intersection, highlighting that bridging structured and unstructured information spaces without losing constraints remains a major open challenge for agentic systems. Code and datasets are publicly available at GitHub (https://github.com/Snowflake-AI-Research/HybridDeepResearch) and Hugging Face (https://huggingface.co/datasets/Snowflake/HybridDeepResearch).