发表机构
Torrens University Australia(澳大利亚托伦斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ESQ-Bench是一款以Oracle为核心的多层级企业NL2SQL基准,实验发现闭源模型在该基准上表现优于开源模型,且隐性语义分歧在高复杂度层级中占比极高。
AI 中文摘要
当前最先进的自然语言转SQL(NL2SQL)模型在Spider和BIRD等已建立的基准上报告执行准确率超过89%。然而,这些基准依赖简化的学术模式和开源SQL方言,无法反映企业数据库环境的复杂性。我们推出ESQ-Bench,这是一款以Oracle为核心的NL2SQL基准,具有系统的复杂性层级,并针对三个企业模式复杂性层级开展隐性分歧评估。我们构建并发布了六个填充模式(465张表、164682行数据、无空表),在Oracle、PostgreSQL、MySQL和SQL Server上使用相同的种子数据,配备包含EM、EX、SR、SD四个指标的评估工具,以及550个经黄金验证的问题-查询对(层级1:95个;层级2:228个;层级3:227个)。使用GPT-4o进行模式关联提示显示,执行匹配度随层级单调下降:2026年6月执行查询的EX分别为79.8%、60.3%和57.2%,而早期142个问题的试点切片则分别为75.6%、80.4%和95.8%。全层级EM均低于7%;在EX通过的查询中,操作层面的隐性分歧达到73%至99%。失败分析显示,更高层级中错误结果语义占主导。Claude Sonnet 4.6使用模式关联提示达到87.4%、74.9%和68.7%的EX(执行查询),在每个层级均超过GPT-4o的模式关联表现。GPT-4o零样本在执行查询上的EX(78.7%、73.5%和77.8%)在层级2至3中与模式关联表现相反,原因是零样本与模式关联分析相比执行率更低且存在幸存者偏差。本地Llama 3.2的模式关联EX仅为全库13.3%(550个中的73个),凸显了闭源API模型与开源权重基线在企业Oracle模式上的差距。
英文摘要
State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments. We introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity tiers. We constructed and released six populated schemas (465 tables, 164,682 rows, zero empty tables) with identical seed data on Oracle, PostgreSQL, MySQL, and SQL Server, a four-metric evaluation harness (EM, EX, SR, SD), and 550 gold-validated question-query pairs (Tier-1: 95; Tier-2: 228; Tier-3: 227). Schema-linked prompting with GPT-4o shows monotonic execution-match degradation across tiers: 79.8, 60.3, and 57.2 percent EX on executed queries (June 2026), versus 75.6, 80.4, and 95.8 percent on an earlier 142-question pilot slice. EM stays below 7 percent tier-wide; operational silent-divergence reaches 73 to 99 percent among EX-passing queries. Failure analysis shows wrong-result semantics dominate at higher tiers. Claude Sonnet 4.6 with schema-linked prompts reaches 87.4, 74.9, and 68.7 percent EX (executed queries), exceeding GPT-4o schema-linked on every tier. GPT-4o zero-shot EX on executed queries (78.7, 73.5, and 77.8 percent) inverts schema-linked at Tiers 2 to 3 due to lower execution rates and survivor bias in the zero-shot versus schema-linked analysis. Local Llama 3.2 schema-linked reaches only 13.3 percent bank-wide EX (73 out of 550), underscoring the gap between closed API models and open-weight baselines on enterprise Oracle schemas.
Comments20 pages, 3 figures, 10 tables