发表机构
University of Texas at Austin; DevRev(德克萨斯大学奥斯汀分校; DevRev)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有NL-to-SQL基准未覆盖的企业模式嵌套结构问题,构建新基准并提出成本感知智能体架构,在新基准上正确率大幅领先,在公开数据集上性能与领先系统相当。
AI 中文摘要
自然语言转SQL系统在学术基准上已快速发展,但生产级企业模式呈现出类似图的半结构化深度嵌套结构,这是当前基准未覆盖的。本文做出两项互补贡献:其一,推出DevRev NL2SQL基准,含900个经执行验证的嵌套类型与链接图结构查询,以及与模式无关的分析推理深度评分标准SDS;其二,提出成本感知单生成智能体架构,其模式选择、元数据检索与错误修复组件均针对该场景需求设计。在DevRev NL2SQL基准上,该系统答案正确率达91.7%,较次优基线高出54.6个百分点;在Spider 2.0 Snowflake公开数据集上,其单生成操作点性能与领先系统相当。
英文摘要
Natural-language-to-SQL systems have ad- vanced rapidly on academic benchmarks, yet production enterprise schemas exhibit graph- like, semi-structured, deeply nested structure that current benchmarks do not measure. We make two complementary contributions. First, we introduce the DevRev NL2SQL bench- mark: 900 execution-verified queries with nested-type and link-graph structure, accom- panied by the Semantic Depth Score (SDS), a schema-agnostic rubric for analytical reasoning depth. Second, we present a cost-aware single- generation agentic architecture whose schema- selection, metadata-retrieval, and error-repair components are designed for the requirements this regime imposes. On the DevRev NL2SQL benchmark the system attains 91.7% answer correctness, a margin of 54.6 percentage points over the next-best baseline; on the Spider 2.0 Snowflake public dataset, it is competitive with leading systems at a single-generation operating point.
Comments17 pages, 3 figures