发表机构
Polytechnique Montréal(蒙特利尔综合理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SQLMorph通过查询变异和细粒度指标,自动生成评估集并揭示系统在复杂查询和语言变化下的性能差异,提升评估的可复现性与实用性。
AI 中文摘要
文本到SQL系统将自然语言查询转换为可执行的SQL,从而实现了对结构化数据的民主化访问。尽管由大型语言模型(LLMs)驱动的近期进展显著,但评估仍然是主要瓶颈:公开基准无法捕捉企业模式的复杂性,而构建私有评估集成本高昂且具有不确定性,使得评估结果难以复现。为解决此问题,我们提出了SQLMorph,一个通过查询变异进行文本到SQL评估的框架。SQLMorph引入了两种自动生成和扩展评估集的技术:连接查询扩展(JQE),通过有效的连接添加系统地增加结构复杂性;以及文本查询增强(TQA),生成受控的自然语言扰动以评估对语言变化的鲁棒性。JQE和TQA创建了针对性的瓶颈点,以挑战特定系统组件。当应用于最先进的系统时,JQE增加了查询覆盖率,并揭示了随着连接数量的增加而出现的准确率下降。同时,TQA表明,由重度缩写引起的语言脆弱性可将准确率降低高达17%。除了评估集之外,SQLMorph还引入了一系列执行级指标,以解决当前二元度量(如执行准确率)的局限性。我们定义了执行精确率(EXP)和执行召回率(EXR),分别量化正确结果和恢复结果的比例,并通过F1分数进行统一评分。我们的实验表明,这些宽松的指标能够对过度预测和欠预测进行细粒度分析,揭示出二元指标所掩盖的系统间差异。综上所述,SQLMorph的查询变异和细粒度指标支持调试,并使文本到SQL评估实践更贴近实际部署。
英文摘要
Text-to-SQL systems translate natural language queries into executable SQL, democratizing access to structured data. Despite recent advances driven by large language models (LLMs), evaluation remains a major bottleneck: public benchmarks fail to capture the complexity of enterprise schema, while building private evaluation sets is costly and nondeterministic, making evaluation results difficult to reproduce. To address this issue, we present SQLMorph, a framework for Text-to-SQL evaluation via query mutation. SQLMorph introduces two techniques to automatically generate and expand evaluation sets: Join Query Expansion (JQE), which systematically increases structural complexity through valid join additions, and Textual Query Augmentation (TQA), which generates controlled natural language perturbations to assess robustness to linguistic variation. JQE and TQA create targeted choke points to challenge specific system components. When applied to state-of-the-art systems, JQE increases query coverage and reveals accuracy degradation as the number of joins grows. Meanwhile, TQA shows that linguistic brittleness induced by heavy abbreviation can reduce accuracy by up to 17%. Beyond evaluation sets, SQLMorph introduces a family of execution-level metrics that address the limitations of current binary measures, such as Execution Accuracy. We define Execution Precision (EXP) and Execution Recall (EXR) to quantify the fraction of correct and recovered results, respectively, and combine them via F1 for unified scoring. Our experiments show that these relaxed metrics enable fine-grained analysis of over- and under-prediction, revealing differences across systems that binary metrics obscure. Together, SQLMorph's query mutation and fine-grained metrics support debugging and better align Text-to-SQL evaluation practices with real-world deployments.
Journal ref2026 IEEE 42nd International Conference on Data Engineering (ICDE), pp. 2628-2640
DOI:10.1109/ICDE65706.2026.00196