arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22830cs.AIcs.LG

超越框架:面向企业文本转SQL的上下文构件端到端优化

Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL

Kate Gwimm, Carson Eisenach

首次发表
浏览论文内容

中文总结 AI 辅助

针对企业文本转SQL任务的上下文瓶颈,本文提出从历史查询生成可复用SQL参考卡片的端到端优化方法,在生产查询基准上取得显著优于优化检索框架的提升,在BEAVER基准上也实现了有效改进。

中文摘要 AI 辅助

将大型语言模型(LLM)部署到企业文本转SQL任务中,瓶颈不在于模型本身,而在于输入模型的上下文:企业业务逻辑涉及数千张表,没有任何模型能一次性读取完整的表目录。因此,本文认为最有效的干预点是模型所使用的知识库上下文,该上下文应从历史使用情况构建,而非作为固定输入进行调优。本文采用查询有向无环图(query-DAG)分解方法——这类中间表示与BEAVER等企业基准所标注的中间表示属于同一类别,此处从生产环境SQL中提取该分解结果,对比了神谕查询图与检索得到的知识库上下文的价值。在该 ablation 实验中,将检索得到的知识库上下文添加到完整神谕图中时,能带来最大的边际提升。在此基础上,本文优化了一种蒸馏流程,将历史查询概要转化为可复用的SQL参考卡片。在某大型在线零售商的5176条生产查询基准上,优化这些上下文构件带来的提升(AST相似度约12%至25%)大于优化检索框架带来的提升(约3%至12%)。在公开BEAVER基准上(该基准缺少本文内部设置中可用的生产使用信号),结果更为复杂:仅表卡片的表现与原始历史SQL相当。最优的优化变体同时检索表卡片和原始SQL,在保留的N=300子集上取得9.00%的分数,而可比基线的分数为6.33%(p值0.12),该优化使用了检索上下文和框架变更,未采用智能体循环。

英文摘要

Deploying LLMs for enterprise Text-to-SQL is bottlenecked less by the model than by what context reaches it: business logic spans thousands of tables, and no model can ingest a full catalog at once. We argue that the most effective place to intervene is therefore the \emph{knowledge-base context} the model consumes, and that this context should be \emph{constructed} from historical usage rather than tuned for as a fixed input. Using a query-DAG decomposition--the same family of intermediates that enterprise benchmarks like BEAVER annotate, here recovered from production SQL--we compare the value of oracle query graphs versus retrieved knowledge-base context. In this ablation, retrieved knowledge-base context provides the largest marginal improvement when added to the full oracle graph. Building on this, we optimize a distillation procedure that turns historical query profiles into reusable SQL reference cards. On a benchmark of 5176 production queries from a major online retailer, optimizing these context artifacts yields larger gains (${\sim}12$--$25\%$ AST similarity) than optimizing the retrieval harness (${\sim}3$--$12\%$). On the public BEAVER benchmark, which lacks the production-usage signals available in our internal setting, the picture is more mixed: table cards alone perform about the same as raw historical SQL. The best optimized variant retrieves both cards and raw SQL, scoring $9.00\%$ versus $6.33\%$ (p-value $0.12$) for the comparable baseline on a held-out $N{=}300$ subset, using retrieved context and harness changes but no agentic loop.

发表机构

  • Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑