发表机构
New Jersey Institute of Technology; Virginia Tech(新泽西理工学院; 弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DexterSQL是基于提示的非微调Text-to-SQL系统,通过深度模式探索器、数据库无关规则创建器、多路径SQL生成三个组件,在BIRD-Dev数据集上提升了开源与闭源大语言模型的Text-to-SQL生成准确率。
AI 中文摘要
基于提示(即非微调)的Text-to-SQL方法,其中底层大语言模型参数未针对该任务进行调整,面临三个问题:(i)依赖粗粒度模式信息,可能无法揭示区分歧义列所需的细粒度关系;(ii)无法捕捉反复出现的SQL生成失败;(iii)在复杂问题中存在条件遗漏、幻觉或错位。本文开发了基于提示/非微调的Text-to-SQL系统DexterSQL,通过三个新颖组件改进SQL生成:(i)深度模式探索器,识别歧义列,分析其单独和联合数据分布以揭示关系及各自的独特作用;(ii)数据库无关规则创建器,仅在训练数据库上挖掘生成SQL与黄金SQL之间的不匹配,并将其转换为数据库无关的修正规则,以捕捉反复出现的LLM失败模式;(iii)多路径SQL生成,引入基于依赖树的中间表示,利用问题的句子结构指导其分解为SQL骨架,用于最终SQL生成。DexterSQL相较于现有最优方法(SOTA)使用开源权重模型和闭源权重模型均实现了更高的准确率。特别是,使用开源权重模型GPT-OSS-120B在BIRD-Dev数据集上,DexterSQL的准确率至少提升2.7%,总准确率达67.6%;使用闭源权重模型时提升幅度更大,使用GPT-4o和GPT-5.2在BIRD-Dev上的总准确率分别为71.6%和72.2%。
英文摘要
Prompting-based (i.e., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not changed for the task, face three problems: (i) relying on coarse-grained schema information that may not reveal the fine-grained relationships needed to distinguish ambiguous columns, (ii) failing to capture recurring SQL-generation failures, and (iii) suffering from omission or hallucination of components in complex questions. This paper develops DexterSQL, a prompting/non-fine-tuning-based Text-to-SQL system that improves SQL generation with three novel components: (i) deep schema explorator that identifies ambiguous columns, analyzes their individual and joint data distributions to uncover their relationships and the distinct role of each, (ii) database-agnostic rule creator that mines mismatches between generated and gold SQL only on the training database and converts them into database-agnostic corrective rules that capture recurring LLM failure patterns; and (iii) multi-path SQL generation that introduces a dependency-tree-based intermediate representation that uses the question's sentence structure to guide its decomposition into an SQL skeleton for final SQL generation. DexterSQL achieves a higher accuracy compared to the state-of-the-art using both open-source/weight and closed-source/weight models. Particularly, DexterSQL shows a high improvement of at least 5.5% using an open-weight model (GPT-OSS-120B) on BIRDDev, with total accuracy 70.4%. DexterSQL also shows better improvement of at least 1.4% using closed-weight models, with total accuracy 72.1% and 72.9% on BIRD-Dev with GPT-4o and GPT-5.2.
CommentsThis version of the paper improved the SQL generation algorithm, increasing the system's overall accuracy. For details, please see the paper