发表机构
College of Medicine and Public Health, Flinders University(弗林德斯大学医学与公共卫生学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对企业级NL2SQL在模式图扩展下的失效问题,提出DRL中间件层,通过动态剪枝等技术限制上下文并验证安全性,在数据库上评估后发现基准代码缺陷的影响,将NL2SQL重新定义为系统工程而非排行榜任务。
AI 中文摘要
在企业级OLTP目录上部署自然语言(NL)接口在规模扩展时失效,原因是语义解析器在模式图扩展下崩溃,导致上下文超出稳定大语言模型(LLM)的注意力预算。我们提出DRL(Deterministic Relational Middleware Layer,确定性关系中间件层),这是一种介于前端与SQL后端之间的安全流水线。DRL包含动态上下文剪枝、关系抽象语法树(AST)类型化以及事务性安全验证(EXPLAIN门控与NULL保护),以限制上下文并标记操作静默分歧(SDop)。我们在PostgreSQL和MySQL上对DRL进行评估,贡献包括:(i)OLTP模式图扩展模型;(ii)包含1000对的工作负载验证套件;(iii)基准B0-B3;(iv)企业级NL2SQL失败分类。在PostgreSQL上,基于模式的提示(B1)相较于朴素的全目录提示(B0)实现76%的上下文减少;DRL的动态路由器(B2)在剪枝p95=0.58ms、中间件p95=4.6ms时达到92%的减少。在修正后的评估工具下,GPT-4o、Claude Sonnet 4.5和Gemini 2.5 Flash的执行匹配率分别为52.9%、52.8%和52.1%;SDop标记了89%-100%的假阳性EX通过查询。GPT-4o的失败主要源于语义/过滤错误(254/471),而列幻觉是次要因素(47/471)。关键的是,我们的评估后处理器中的单个正则表达式缺陷静默地抑制了准确率,并产生了4%-10%的虚假跨供应商差距,该差距在修正后消失,表明基准代码应与其评分的模型受到同等审查。DRL将企业级NL2SQL重新定义为系统工程——包括上下文限制、验证和感知计划的准入——而非排行榜练习。
英文摘要
Deploying natural-language interfaces over enterprise OLTP catalogs fails at scale because semantic parsers collapse under schema-graph scaling, inflating context beyond stable LLM attention budgets. We present DRL (Deterministic Relational Middleware Layer), a safe pipeline interposing between front-ends and SQL backends. DRL comprises dynamic context pruning, relational AST typing, and transactional safeguard verification (EXPLAIN gating and NULL guards) to bound context and flag operational silent divergence (SDop). We evaluate DRL on PostgreSQL and MySQL, contributing (i) an OLTP schema-graph scaling model, (ii) a 1,000-pair Workload Verification Suite, (iii) baselines B0-B3, and (iv) an enterprise NL2SQL failure taxonomy. On PostgreSQL, schema-linked hints (B1) yield a 76% context reduction over naive full-catalog prompting (B0); DRL's dynamic router (B2) reaches a 92% reduction at pruning p95 = 0.58 ms and middleware p95 = 4.6 ms. GPT-4o, Claude Sonnet 4.5, and Gemini 2.5 Flash achieve 52.9%, 52.8%, and 52.1% execution match under a corrected evaluation harness; SDop flags 89-100% of false-positive EX-passing queries. GPT-4o failures are dominated by semantic/filter errors (254/471), while column hallucination is a minor factor (47/471). Crucially, a single regex defect in our evaluation post-processor silently suppressed accuracy and manufactured a false 4-10% cross-vendor gap that vanished when corrected, showing that benchmark code deserves the same scrutiny as the models it scores. DRL reframes enterprise NL2SQL as systems engineering - context bounding, verification, and plan-aware admission - not a leaderboard exercise.
Comments34 pages, 4 figures, 10 tables. Includes 1