arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReToolSQL:用于鲁棒文本到SQL的智能体强化学习

ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL

Pratik Kakkar, Chandra Dhir, Ravi Shankar, Pareekshit Reddy Gaddam, Anup Shirgaonkar

arXiv 2608.27796首次发表:更新:

发表机构

JPMorganChase(摩根大通)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReToolSQL为文本到SQL提出两阶段训练框架,结合监督微调与智能体强化微调,在BIRD-SQL基准上取得高执行准确率,登顶单模型开发集排行榜,是实现企业级鲁棒文本到SQL的可行方案。

AI 中文摘要

近期研究表明,基于执行反馈的强化学习可大幅提升文本到SQL(Text-to-SQL)的性能,通常能让较小规模的模型达到或超过大得多的系统的表现。然而,现有大多数方法将SQL生成为单轮任务,限制了模型通过迭代细化从错误中恢复的能力。本文提出ReToolSQL,一种用于文本到SQL的两阶段训练框架,结合了:(i)对拒绝采样的推理轨迹进行监督式预热,以及(ii)对多轮工具使用轨迹进行智能体强化微调(Reinforcement Fine-Tuning, RFT)。核心洞见在于,两个阶段作用于互补维度:在经验证的特权教师轨迹上的监督微调(Supervised Fine-Tuning, SFT)扩大了可解决问题的集合(提升了最难案例的pass@k覆盖率);而RFT通过教导模型何时验证、检索何种证据以及如何根据执行反馈修复有缺陷的SQL,将这种扩大的能力转化为更高的单轮准确率。应用于指令微调后的Gemma 4(31B参数),仅RFT在BIRD-SQL开发基准上达到73.66%的执行准确率(Execution Accuracy, EX)(结合自一致性后为74.12% EX)。从SFT检查点初始化RFT(SFT→RFT)得到我们最强的模型,单轮执行准确率达74.32%,结合自一致性后为74.77% EX。截至撰写本文时,该模型在BIRD单模型开发集排行榜上排名第一。该方法使用基于执行正确性的复合奖励,除基准本身外无需人工标注,且在单一31B参数的密集模型内运行,表明精心设计的、基于工具使用轨迹的SFT→RFT流程是实现企业级鲁棒文本到SQL的可行路径。

英文摘要

Recent work has shown that reinforcement learning from execution feedback can substantially improve text-to-SQL performance, often enabling smaller models to match or exceed much larger systems. However, most existing approaches treat SQL generation as a single-turn task, limiting the model's ability to recover from errors through iterative refinement. We present ReToolSQL, a two-stage training framework for text-to-SQL that combines (i) a supervised warm-start on rejection-sampled reasoning traces with (ii) agentic reinforcement fine-tuning (RFT) over multi-turn tool-use trajectories. The key insight is that the two stages act on complementary axes, the supervised fine-tuning (SFT) on verified privileged-teacher traces expands the set of solvable questions (raising pass@k coverage on the hardest cases), while RFT converts that expanded capability into higher single-pass accuracy by teaching the model when to verify, what evidence to retrieve, and how to repair faulty SQL from execution feedback. Applied to Gemma 4 instruction-tuned (31B), RFT alone achieves 73.66% execution accuracy (EX) on the BIRD-SQL development benchmark (74.12% EX with self-consistency). Initializing RFT from the SFT checkpoint (SFT$\to$RFT) yields our strongest model at 74.32% EX single-pass and 74.77% EX with self-consistency. At the time of writing, this ranked first on the BIRD single-model development-set leaderboard. The approach uses composite rewards anchored on execution correctness, requires no human annotation beyond the benchmark itself, and operates within a single dense 31B model, showing that a properly designed SFT$\to$RFT pipeline over tool-use trajectories is a practical path toward robust enterprise-grade text-to-SQL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑