arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ACTS-SQL:基于大型语言模型的面向智能体与评判器的树结构化SQL正确性校验

ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models

Xinmei Huang, Jie Song, Peng Li, Fuxin Jiang, Jing Zhang, Tieying Zhang, Jianjun Chen, Chenming Liu, Tao Yang, Maoyin Liu, Wenda Li, Hong Chen, Cuiping Li

arXiv 2608.15145首次发表:更新:

发表机构

Renmin University of China; ByteDance Inc.(中国人民大学; 字节跳动公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出无训练框架ACTS-SQL,将SQL修正建模为计划引导的树结构化调试过程,集成执行校验与诊断工具,在BIRD-Critic基准及火山引擎TLS生产环境中均实现显著准确率提升。

AI 中文摘要

大型语言模型(LLMs)已越来越多地被应用于Text-to-SQL系统,但SQL错误仍是实际Text-to-SQL推理流程中的主要障碍。现有SQL修正方法要么依赖大规模高质量训练数据,开销巨大;要么采用单路径智能体工作流,易受早期错误影响且存在错误传播问题。为开发适用于工业场景的实用SQL正确性系统,本文提出一种无训练框架,将SQL修正建模为计划引导的树结构化调试过程,通过维护多种修正策略并支持回溯,缓解迭代优化过程中的错误累积。我们还集成了基于执行的校验与子句级诊断工具,以支持策略剪枝与精确错误定位。在BIRD-Critic基准上对系统进行评估,结果显示其在强大的LLM主干模型与代表性智能体基线方法上均实现了持续的准确率提升,较此前的最优方法提升了9.42%。该框架还已部署至火山引擎的Torch Log Service(TLS),以支持在线Text-to-TLS API。在生产环境中,采用代表性强LLM主干模型(GPT-5)时,其在真实用户查询上将执行准确率从36.77%提升至53.61%。这些结果证明了本文方法在实际部署中的有效性与稳定性。

英文摘要

Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines. Existing SQL correction approaches either rely on large-scale, high-quality training data with substantial overhead, or adopt single-path agentic workflows that are brittle to early mistakes and prone to error propagation. To develop a practical SQL correctness system for industrial scenarios, we present a training-free framework that formulates SQL correction as a plan-guided, tree-structured debugging process. By maintaining multiple correction strategies and enabling backtracking, the framework mitigates error accumulation during iterative refinement. We further integrate execution-based verification and clause-level diagnostic tools to support strategy pruning and precise error localization. We evaluate the system on the BIRD-Critic benchmark and observe consistent accuracy gains over strong LLM backbones and representative agent-based baselines, achieving a 9.42% improvement over the previous state-of-the-art method. The framework is also deployed in the Torch Log Service (TLS) of Volcano Engine to support an online Text-to-TLS API. In production, it improves execution accuracy from 36.77% to 53.61% on real user queries with a representative strong LLM backbone (GPT-5). These results demonstrate the effectiveness and stability of our approach in real-world deployments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑