COAL-SQL:面向文本到SQL后训练的覆盖引导增强与失败驱动学习
COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training
浏览论文内容
中文总结 AI 辅助
COAL-SQL通过覆盖引导增强和失败驱动学习,以少量后训练数据提升文本到SQL生成的结构覆盖与执行准确率。
中文摘要 AI 辅助
文本到SQL任务将自然语言问题转换为可执行的SQL查询,但开源大语言模型在面向复杂真实世界SQL生成时,仍需针对特定任务进行后训练。有效的后训练既需要覆盖目标任务所需能力的训练数据,也需要能使模型习得这些能力的学习策略。现有数据集提供了有价值的监督信号,但对SQL结构的覆盖并不完整,而增强方法通常只是扩充数据,并未识别结构缺口。此外,单独使用监督微调(SFT)或强化学习(RL)无法动态应对训练中暴露的弱点。我们提出了COAL-SQL,一个结合覆盖引导增强(CGA)和失败驱动学习(FDL)的统一框架。CGA利用贪心选择识别原始数据集中缺失的SQL结构,并构建互补示例,从而提升结构覆盖率。FDL保留GRPO作为主要优化目标,同时为未解决的示例提供有针对性的监督。在步骤层面,它对由强LLM生成的已验证推理轨迹应用SFT,以处理累积的失败。在轮次层面,它基于累积失败检索结构相关的示例,以创建有针对性的练习,帮助模型习得相应的SQL能力。仅使用12,600个不同的后训练示例,COAL-SQL在BIRD开发集上达到了64.9%的执行准确率,并优于在相当规模下训练的基线模型。代码可在该https URL获取。
英文摘要
Text-to-SQL translates natural-language questions into executable SQL queries, but open-source large language models still require task-specific post-training for complex, real-world SQL generation. Effective post-training requires both training data that cover the capabilities demanded by the target task and a learning strategy that enables the model to acquire them. Existing datasets provide valuable supervision but incompletely cover SQL structures, while augmentation methods typically expand data without identifying structural gaps. Moreover, supervised fine-tuning (SFT) or reinforcement learning (RL) alone cannot dynamically address weaknesses exposed during training. We propose COAL-SQL, a unified framework combining Coverage-Guided Augmentation (CGA) and Failure-Driven Learning (FDL). CGA uses greedy selection to identify SQL structures missing from the original dataset and constructs complementary examples, improving structural coverage. FDL retains GRPO as the main optimization objective while supplying targeted supervision for unsolved examples. At the step level, it applies SFT to verified reasoning traces generated by a strong LLM for accumulated failures. At the epoch level, it retrieves structurally related examples based on accumulated failures to create targeted practice, helping the model acquire the corresponding SQL capabilities. With only 12,600 distinct post-training examples, COAL-SQL achieves 64.9% execution accuracy on the BIRD development set and outperforms baselines trained at comparable scale. The code is available at https://github.com/TechNomad-ds/COAL-SQL.
发表机构
- Peking University(北京大学)
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。