arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向结构化数据搜索的引导式表格检索

Guided Table Retrieval for Structured Data Search

Alekh Jindal, Jyoti Pandey, Christina Pavlopoulou, Ronith PR, Sharath Prakash, Shi Qiao, Shivani Tripathi, Wangda Zhang

arXiv 2608.11644首次发表:更新:

AI 中文总结

该研究针对结构化数据库自然语言问答任务,提出四阶段引导式表格检索流程,在BIRD-DEV和BEAVER基准上取得优于基线的准确率与F1值,生成的精确连接树可直接供下游查询编译器使用。

AI 中文摘要

针对结构化数据库回答自然语言问题,需识别相关表格并确定如何连接它们,该任务既需要模式知识,也需要对用户意图的语义理解。本文提出引导式表格检索,这是一个四阶段流程,结合基于哈希预测器的确定性定位、连接图可达性的结构探索、大语言模型(LLM)驱动的源与目标消歧,以及算法合并为最小的拓扑有序连接树。通过将问题分解为具有不同职责的阶段——确定性、覆盖性、语义推理和连贯性,该流程避免了端到端LLM方法的脆弱性,同时在最需要上下文判断的地方利用LLM。我们在BIRD-DEV和企业级BEAVER基准上进行评估,分别达到94%和70%的准确率,92%和53%的F1值,在准确率和F1上显著优于现有基线,且生成可直接供下游查询编译器使用的精确连接树。

英文摘要

Answering natural language questions over structured databases requires identifying the relevant tables and determining how to join them---a task that demands both schema knowledge and semantic understanding of the user's intent. We present guided table retrieval, a four-phase pipeline that combines deterministic grounding via hash-based predictors, structural exploration of join-graph reachability, LLM-powered disambiguation of sources and targets, and algorithmic merging into minimal, topologically ordered join trees. By decomposing the problem into phases with distinct responsibilities--- determinism, coverage, semantic reasoning, and coherence--- the pipeline avoids the brittleness of end-to-end LLM approaches while leveraging LLMs where their contextual judgment is most needed. We evaluate on BIRD-DEV and the enterprise-scale BEAVER benchmark, achieving 94% and 70% precision respectively, with 92% and 53% F1---substantially outperforming existing baselines on precision and F1 while producing exact join trees that can be directly consumed by downstream query compilers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑