AI 中文总结
提出语义路径编译(SPC)系统,在ACME保险基准测试中,其文本到SQL任务正确率达97.4%,显著优于基线系统,且鲁棒性更强。
AI 中文摘要
直接文本到SQL任务要求语言模型完成两项工作:解读业务问题并构建完整的关系型查询。在企业级模式中,SQL可能执行成功,但使用了错误的关系角色或聚合粒度。本研究探讨了随机边界的替代放置方式:多轮规划器对短语进行 grounding(落地),并从问题特定的受控选项中进行选择;图遍历、角色谓词、粒度降低、SQL构建及确定性检查均通过代码实现。我们在ACME保险基准测试中,将该语义路径编译(SPC)系统与直接DDL到SQL生成进行对比评估。在包含38个问题的经裁决的对比集(每个问题3次运行)中,SPC在37个问题的所有运行中均被裁决为正确,占比97.4%,而基线系统的这一比例为55.3%(21个问题)。配对不一致结果显示,16个问题倾向于SPC,无问题倾向于基线,双侧精确McNemar检验p值为3.05×10^-5。SPC在114次运行结果中,至少有一次正确回答了全部38个问题,产生1次弃权(不执行),无裁决为错误但已执行的运行;而基线系统在同一组中产生29次裁决错误运行,另有7次额外的、由评判者标记的仅数据巧合的情况。严格等价性敏感性分析扩大了配对差异。使用GPT-5.4和Gemini-3.6-Flash进行的额外SPC运行显示出相似的问题级鲁棒性,尽管它们的每次运行裁决异常未保留。全项目分析中保留了6个额外基准项目,并按失败类别单独记录。该研究支持端到端系统层面的结果,而非仅编译本身带来性能提升的因果主张,因为SPC接收了DDL基线未具备的受控语义人工制品。
英文摘要
Direct text-to-SQL asks a language model to do two jobs: interpret the business question and construct the complete relational query. In enterprise schemas, SQL can execute successfully while using the wrong relationship role or aggregation grain. We study an alternative placement of the stochastic boundary. A multi-turn planner grounds phrases and selects from question-specific governed options; graph traversal, role predicates, grain lowering, SQL construction, and deterministic checks are implemented in code. We evaluate this semantic path compilation (SPC) system against direct DDL-to-SQL generation on the ACME insurance benchmark. On a 38-question adjudicated comparison set with three runs per question, SPC was adjudicated correct on every run for 37 questions (97.4%), compared with 21 (55.3%) for the baseline. The paired discordance was 16 questions in favor of SPC and none in favor of the baseline (two-sided exact McNemar p=3.05x10^-5). SPC answered all 38 questions correctly at least once and produced one refusal and no adjudicated wrong-but-executed run across 114 run outcomes; the baseline produced 29 adjudicated wrong runs and seven additional judge-flagged data-only coincidences on the same set. A strict-equivalence sensitivity analysis increased the paired difference. Additional SPC runs with GPT-5.4 and Gemini-3.6-Flash showed similar question-level robustness, although their per-run verdict artifacts were not preserved. Six additional benchmark items are retained in an all-item analysis and documented separately by failure class. The study supports an end-to-end systems result, not a causal claim that compilation alone produced the gain, because SPC receives governed semantic artifacts that the DDL baseline does not.
Comments10 sections, 2 figures, 6 tables. Preprint. Code and research artifacts are described in the manuscript