arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Baikal:面向数据湖深度研究的结构化搜索框架

Baikal: Structured Search for Deep Research over Data Lakes

Dhruv Agarwal, Rishitha Guttapalle Mohan, Aarti Kumari, Ashi Sinha, Athulya Anil, Kavitha Srinivas, Horst Samulowitz, Andrew McCallum

arXiv 2607.27726首次发表:更新:

AI 中文总结

Baikal是面向数据湖深度研究的结构化搜索框架,通过聚类证据为语义区域并自适应搜索,在HybridQA和TAT-QA数据湖上较基线方法提升报告得分28%和36%,展现了结构化语义探索的价值。

AI 中文摘要

面向数据湖的深度研究需要大语言模型(LLM)智能体调查数千个异构表格与段落中的证据以生成综合报告。现有方法采用迭代检索与生成模式,让累积上下文决定后续调查方向,这可能过度利用局部有前景的证据,且在固定预算下无法覆盖不同语义区域。为解决该问题,我们将数据湖深度研究视为带预算的搜索问题,提出Baikal框架:该框架将异构证据聚类为语义区域,再在区域间自适应搜索以平衡探索与利用。在每个选定区域内,Baikal生成并调查基于区域的子问题,以发现质量作为奖励更新区域级价值估计,并在从随机选择、LLM引导选择到贝叶斯ε-贪心算法、UCB算法的各类策略下指导搜索。我们在两个数据湖上对Baikal进行评估:HybridQA数据湖含10993个表格与22.7万条维基百科段落,TAT-QA数据湖含2757个表格与1.3万条财务报告段落,每个数据湖各用15个查询。我们采用涵盖基于性、相关性、多样性、实用性的新评分标准,用GPT-5-mini对Baikal及DeepSearcher、带检索与聚类变体的OpenCode研究智能体等强基线评分。在两个数据湖上,Baikal在多种区域选择策略下表现优异;其最优配置在HybridQA上将报告得分较最强基线提升28%,在TAT-QA上提升36%。我们的分析将这些提升归因于组织并探索语义证据区域,这在相同子问题预算下改善了基于性与多样性,产出了更有用的发现。这些结果证明了结构化语义探索对异构数据湖的系统研究与发现的价值。

英文摘要

Deep research over data lakes requires an LLM agent to investigate evidence across thousands of heterogeneous tables and passages to synthesize a report. Existing methods perform iterative retrieval and generation, letting accumulated context determine what to investigate next, which can overexploit locally promising evidence and fail to cover distinct semantic regions under a fixed budget. To address this, we cast deep research over data lakes as a budgeted search problem and present Baikal - a framework that clusters heterogeneous evidence into semantic regions, then searches over them adaptively to balance exploration and exploitation. Within each selected region, Baikal generates and investigates region-grounded subquestions, using finding quality as rewards to update region-level value estimates and guide search under policies ranging from random and LLM-guided selection to Bayesian $ε$-greedy and UCB. We evaluate Baikal on 15 queries each over HybridQA and TAT-QA data lakes containing 10,993 and 2,757 tables, respectively, together with 227K Wikipedia passages and 13K financial report passages. We assess research quality with a new rubric covering groundedness, relevance, diversity, and utility, and use GPT-5-mini to score Baikal and strong baselines, including DeepSearcher and an OpenCode research agent with retrieval and clustering variants. Across both data lakes, Baikal performs strongly under several region-selection policies; its best configuration improves report scores over the strongest baselines by 28% on HybridQA and 36% on TAT-QA. Our analyses attribute these gains to organizing and exploring semantic evidence regions, which improves groundedness and diversity and yields more useful findings under the same subquestion budget. These results demonstrate the value of structured semantic exploration for systematic research and discovery over heterogeneous data lakes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑