发表机构
Renmin University of China; ByteDance Inc.(中国人民大学; 字节跳动公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无训练框架ACTS-SQL,将SQL修正建模为计划引导的树结构化调试过程,集成执行校验与诊断工具,在BIRD-Critic基准及火山引擎TLS生产环境中均实现显著准确率提升。
AI 中文摘要
大型语言模型(LLMs)已越来越多地被应用于Text-to-SQL系统,但SQL错误仍是实际Text-to-SQL推理流程中的主要障碍。现有SQL修正方法要么依赖大规模高质量训练数据,开销巨大;要么采用单路径智能体工作流,易受早期错误影响且存在错误传播问题。为开发适用于工业场景的实用SQL正确性系统,本文提出一种无训练框架,将SQL修正建模为计划引导的树结构化调试过程,通过维护多种修正策略并支持回溯,缓解迭代优化过程中的错误累积。我们还集成了基于执行的校验与子句级诊断工具,以支持策略剪枝与精确错误定位。在BIRD-Critic基准上对系统进行评估,结果显示其在强大的LLM主干模型与代表性智能体基线方法上均实现了持续的准确率提升,较此前的最优方法提升了9.42%。该框架还已部署至火山引擎的Torch Log Service(TLS),以支持在线Text-to-TLS API。在生产环境中,采用代表性强LLM主干模型(GPT-5)时,其在真实用户查询上将执行准确率从36.77%提升至53.61%。这些结果证明了本文方法在实际部署中的有效性与稳定性。
英文摘要
Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines. Existing SQL correction approaches either rely on large-scale, high-quality training data with substantial overhead, or adopt single-path agentic workflows that are brittle to early mistakes and prone to error propagation. To develop a practical SQL correctness system for industrial scenarios, we present a training-free framework that formulates SQL correction as a plan-guided, tree-structured debugging process. By maintaining multiple correction strategies and enabling backtracking, the framework mitigates error accumulation during iterative refinement. We further integrate execution-based verification and clause-level diagnostic tools to support strategy pruning and precise error localization. We evaluate the system on the BIRD-Critic benchmark and observe consistent accuracy gains over strong LLM backbones and representative agent-based baselines, achieving a 9.42% improvement over the previous state-of-the-art method. The framework is also deployed in the Torch Log Service (TLS) of Volcano Engine to support an online Text-to-TLS API. In production, it improves execution accuracy from 36.77% to 53.61% on real user queries with a representative strong LLM backbone (GPT-5). These results demonstrate the effectiveness and stability of our approach in real-world deployments.