发表机构
Shanghai Jiao Tong University; University of New South Wales; Fudan University; Wuhan University of Technology(上海交通大学; 新南威尔士大学; 复旦大学; 武汉理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究文本到SQL问题,提出EvoSQL协同进化框架,通过生成器与评论家迭代交互、维护内存、验证候选及自蒸馏策略优化微调等,在Spider和BIRD实验中持续改进开源模型,证明该方法是有效途径。
AI 中文摘要
随着大语言模型的发展,文本到SQL取得了快速进展,但复杂数据库查询仍需超出一次性生成的推理,包括多步分解、基于执行的诊断和有针对性的修正。我们提出了EvoSQL,一种将SQL合成表述为生成器和评论家之间迭代交互的协同进化框架。EvoSQL维护上下文候选内存,通过执行信号和基于大语言模型的评论来验证SQL候选,并通过效用引导聚合更新内存。为强化基础的生成器-评论家对,还引入了自蒸馏策略优化微调阶段,将执行感知监督注入现代编码大语言模型骨干。在Spider和BIRD上的实验表明,EvoSQL持续改进开源模型,SDPO初始化进一步提升了选定骨干在特定测试集上的表现。这些结果表明基于内存的协同进化是通往更可靠和通用的文本到SQL系统的有效途径。
英文摘要
Text-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generation, including multi-step decomposition, execution-based diagnosis, and targeted correction. We present EvoSQL, a co-evolution framework that formulates SQL synthesis as an iterative interaction between a generator and a critic. EvoSQL maintains a contextualized candidate memory, verifies SQL candidates with both execution signals and LLM-based critique, and updates its memory through utility-guided aggregation. To strengthen the underlying generator-critic pair, we further introduce a Self-Distillation Policy Optimization (SDPO) fine-tuning stage that injects execution-aware supervision into modern coding LLM backbones. Experiments on Spider and BIRD show that EvoSQL consistently improves open-source models over Maj@16 baselines, with particularly large gains on BIRD-Dev, ranging from +1.37% for Qwen3-4B to +9.19% for Qwen2.5-Coder-3B. SDPO initialization further improves selected backbones on Spider-Test and BIRD-Dev. These results suggest that memory-grounded co-evolution is an effective path toward more reliable and generalizable Text-to-SQL systems. Code is available at https://github.com/valleysprings/EvoSQL.
CommentsAccepted as EMNLP Findings (2026)