AI 中文总结
该研究针对结构化数据库自然语言问答任务,提出四阶段引导式表格检索流程,在BIRD-DEV和BEAVER基准上取得优于基线的准确率与F1值,生成的精确连接树可直接供下游查询编译器使用。
AI 中文摘要
针对结构化数据库回答自然语言问题,需识别相关表格并确定如何连接它们,该任务既需要模式知识,也需要对用户意图的语义理解。本文提出引导式表格检索,这是一个四阶段流程,结合基于哈希预测器的确定性定位、连接图可达性的结构探索、大语言模型(LLM)驱动的源与目标消歧,以及算法合并为最小的拓扑有序连接树。通过将问题分解为具有不同职责的阶段——确定性、覆盖性、语义推理和连贯性,该流程避免了端到端LLM方法的脆弱性,同时在最需要上下文判断的地方利用LLM。我们在BIRD-DEV和企业级BEAVER基准上进行评估,分别达到94%和70%的准确率,92%和53%的F1值,在准确率和F1上显著优于现有基线,且生成可直接供下游查询编译器使用的精确连接树。
英文摘要
Answering natural language questions over structured databases requires identifying the relevant tables and determining how to join them---a task that demands both schema knowledge and semantic understanding of the user's intent. We present guided table retrieval, a four-phase pipeline that combines deterministic grounding via hash-based predictors, structural exploration of join-graph reachability, LLM-powered disambiguation of sources and targets, and algorithmic merging into minimal, topologically ordered join trees. By decomposing the problem into phases with distinct responsibilities--- determinism, coverage, semantic reasoning, and coherence--- the pipeline avoids the brittleness of end-to-end LLM approaches while leveraging LLMs where their contextual judgment is most needed. We evaluate on BIRD-DEV and the enterprise-scale BEAVER benchmark, achieving 94% and 70% precision respectively, with 92% and 53% F1---substantially outperforming existing baselines on precision and F1 while producing exact join trees that can be directly consumed by downstream query compilers.