arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12310cs.CLcs.AI

LakeQuest:用于跨数据湖的有基础问答的三领域基准测试

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

Michael Solodko, Steven Gong, Guangwei Yu, Satya Krishna Gorti, Jesse C. Cresswell, Victor Zhong

AI总结:

介绍LakeQuest基准测试,用于评估跨数据湖的有基础问答。它跨越三个领域,含9846个QA对及证据指针,能暴露系统故障模式。通过基线评估发现高质量检索不能保证正确推理,凸显未来智能QA系统需强大发现和跨文件组合机制。

AI中文摘要:

虽然现代问答(QA)系统在干净、模式对齐的语料库上表现出色,但现实世界中的知识很少如此整齐地打包。在企业和科学数据湖上回答问题需要系统在异构、弱结构化的表、段落和链接的元数据集合中导航。当前的基准测试忽略了这个嘈杂的发现过程,无法评估端到端性能。为了弥合这一差距,我们引入了LakeQuest,这是一个经过人工验证的包含9846个QA对的基准测试,旨在评估在现实数据湖上的端到端检索与合成流程。LakeQuest跨越三个不同领域(人工智能/机器学习元数据、零售银行和多模态生物医学药物信息),并为每个问题配对精确的、模态感知证据指针。通过将源发现与跨模态合成分离,LakeQuest揭示了现代QA系统中的关键故障模式。我们的基线评估,包括标准的检索增强生成(RAG)和智能工具使用方法,表明高质量的检索并不能保证正确的推理。系统在元数据图中的关系链接、银行账本中的政策基础以及生物医学背景下的联合表格QA方面一直存在困难,这突出了未来智能QA系统中强大发现和可靠跨文件组合机制的必要性。

英文摘要:

While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scientific data lakes requires systems to navigate heterogeneous, weakly structured collections of tables, passages, and linked metadata. Current benchmarks abstract away this noisy discovery process, failing to evaluate end-to-end performance. To bridge this gap, we introduce LakeQuest, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes. LakeQuest spans three diverse domains (AI/ML metadata, retail banking, and multimodal biomedical drug information) and pairs every question with exact, modality-aware evidence pointers. By isolating source discovery from cross-modal synthesis, LakeQuest exposes critical failure modes in modern QA systems. Our baseline evaluations, including standard Retrieval-Augmented Generation (RAG) and agentic tool-use methods, reveal that high-quality retrieval does not guarantee correct reasoning. Systems consistently struggle with relation chaining in metadata graphs, policy grounding in bank ledgers, and joint tabular QA in biomedical contexts, highlighting the need for robust discovery and faithful cross-file composition mechanisms in future agentic QA systems.

补充信息

↑