发表机构
University of Manchester; Faculty of Science and Engineering(曼彻斯特大学; 科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对多数据库Text-to-SQL场景提出分层模式链接框架MDB-Link,结合LLM实现数据库定位与列选择,在多个数据集上性能及运行速度均优于基线方法,可高效支撑SQL生成。
AI 中文摘要
传统Text-to-SQL研究及基准均假设目标数据库已知,忽略了查询需在大型异构数据库集合中路由的场景。为此,本文研究多数据库场景下的模式链接问题,该场景中系统需先定位目标数据库,再构建与SQL生成相关的紧凑模式。本文提出MDB-Link,一种分层模式链接框架,其从全局索引中检索与问题相关的列,聚合检索证据以筛选数据库短名单,并使用感知预算的大语言模型(LLM)进行数据库重排序、表选择及列关联。采用Qwen2.5-14B时,MDB-Link在MMQA、Spider2-Snow、BIRD-dev数据集上的数据库定位和列选择性能优于LinkAlign,且生成的模式子集规模接近真实模式;在MMQA上精确匹配度从16.88提升至51.41,Spider2-Snow上从2.50提升至9.17,BIRD-dev上从12.52提升至38.01。此外,MDB-Link的运行速度快于LinkAlign和AutoLink,证明了分层模式缩减对下游SQL生成的有效性。
英文摘要
Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, heterogeneous database collection. We therefore study schema linking in a multi-database setting, where the system must first locate the target database and then construct a compact, SQL-relevant schema for generation. We propose MDB-Link, a hierarchical schema-linking framework that retrieves question-relevant columns from a global index, aggregates retrieval evidence to shortlist databases, and uses a budget-aware large language model (LLM) for database reranking, table selection, and column grounding. With Qwen2.5-14B, MDB-Link outperforms LinkAlign on MMQA, Spider2-Snow, and BIRD-dev in database localization and column selection while producing schema subsets close in size to the gold schemas. Exact match improves from 16.88 to 51.41 on MMQA, 2.50 to 9.17 on Spider2-Snow, and 12.52 to 38.01 on BIRD-dev. MDB-Link also runs faster than LinkAlign and AutoLink, demonstrating the effectiveness of hierarchical schema reduction for downstream SQL generation.