发表机构
Korea Advanced Institute of Science and Technology(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SafeQL是将DBMS设为主动引导者的基于搜索的Text-to-SQL优化方法,可修复SQL错误,在Bird和Spider基准上提升了执行准确率与效率。
AI 中文摘要
大语言模型(LLM)通过提供无需特定任务微调的数据库自然语言接口,推动了Text-to-SQL技术的发展。但现有的LLM系统仍存在不可靠问题,常生成不符合数据库模式的无效SQL查询,引用不存在的表、属性、函数或值。此类错误持续存在,是因为与数据库管理系统(DBMS)的交互通常仅限于错误消息,使其在查询优化过程中基本处于被动角色。本文提出SafeQL,一种基于搜索的优化范式,将DBMS的角色重新定义为优化过程中的主动引导者。SafeQL不在执行失败后重新生成整个查询,而是解释DBMS的反馈,仅逐步修复错误组件。每个优化步骤被表述为在“安全查询空间”内的引导搜索,候选查询通过DBMS执行逐步验证,从而收敛到可执行查询并避免重复生成错误。在Bird和Spider基准上的实验表明,与基于重新生成的方法相比,SafeQL显著提升了执行准确率和效率。
英文摘要
Large language models (LLMs) have advanced Text-to-SQL by enabling natural language interfaces to databases without task-specific fine-tuning. However, existing LLM-based systems remain unreliable, often generating SQL queries that are invalid under the database schema, referencing non-existent tables, attributes, functions, or values. Such errors persist because interactions with the database management system (DBMS) are typically limited to error messages, leaving it in a largely passive role during query refinement. This paper proposes SafeQL, \textit{a search-based refinement paradigm that redefines the role of the DBMS as an active guide in the refinement process}. Instead of regenerating entire queries after execution failure, SafeQL interprets DBMS feedback to incrementally repair only the erroneous components. Each refinement step is formulated as a guided search within a \textit{safe query space}, where candidate queries are progressively validated through DBMS execution, thereby converging to an executable query and preventing repeated regeneration of errors. Experiments on the Bird and Spider benchmarks show that SafeQL significantly improves execution accuracy and efficiency compared to regeneration-based methods.
CommentsVLDB 2026
Journal refProc. VLDB Endow. 19(9): 2210-2223, 2026