arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SERL-SQL:面向Text-to-SQL强化智能体学习的选择性回溯蒸馏

SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning

Tao Liu, Tao Feng, Xiangheng Li, Jinwang Song, Yifan Li, Xiaoqing Cheng, Dixuan Zhang, Siquan Li, Lin Lan, Hongying Zan, Kunli Zhang, Chao Wu

arXiv 2608.00485首次发表:更新:

AI 中文总结

该研究针对现有Text-to-SQL强化学习方法奖励指导不足的问题,提出SERL-SQL框架,通过教师模型重评分与掩码权重优化GRPO优势,在BIRD、Spider等基准上取得优异执行准确率,其奖励选择策略表现优于一致性选择。

AI 中文摘要

当前Text-to-SQL系统越来越依赖多轮交互、执行反馈和强化学习。然而,大多数现有方法仅将执行正确性作为轨迹级奖励,这对识别导致成功或失败的SQL决策的指导作用有限。我们提出SERL-SQL,一种面向多轮Text-to-SQL智能体的基于执行的选择性强化学习框架。SERL-SQL对同策略SQL交互轨迹进行采样,并使用仅训练阶段的教师模型,通过执行反馈对学生动作重新评分。由此产生的教师-学生似然差距被转换为有界的掩码权重,仅在SQL和工具动作令牌上对GRPO优势进行重新加权。通过这种方式,任务奖励保留了优化方向,而执行回溯则提供了本地化的信用分配。在BIRD、Spider和跨域基准上的实验表明,SERL-SQL取得了具有竞争力的性能,在BIRD-Dev上达到76.56%的执行准确率,在Spider-Test上达到89.92%。此外,我们基于奖励的选择策略非常接近最佳N候选的上界,且始终优于基于一致性的选择,这表明SERL-SQL生成的高质量候选可通过轻量级的基于执行的奖励可靠识别。我们的代码将在此httpsURL发布。

英文摘要

Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods use execution correctness only as a trajectory-level reward, which provides limited guidance for identifying the SQL decisions responsible for success or failure. We propose SERL-SQL, a selective execution-grounded reinforcement learning framework for multi-turn Text-to-SQL agents. SERL-SQL samples on-policy SQL interaction trajectories and uses a training-only teacher to re-score student actions with execution feedback. The resulting teacher--student likelihood gap is converted into bounded, masked weights that reweight GRPO advantages only on SQL and tool-action tokens. In this way, task rewards preserve the optimization direction, while execution hindsight provides localized credit assignment. Experiments on BIRD, Spider, and cross-domain benchmarks show that SERL-SQL achieves competitive performance, reaching 76.56% execution accuracy on BIRD-Dev and 89.92% on Spider-Test. Moreover, our reward-based selection strategy closely approaches the oracle Best-of-N upper bound and consistently outperforms consistency-based selection, showing that SERL-SQL produces high-quality candidates that can be reliably identified by lightweight execution-grounded rewards. Our code will be released at https://github.com/Ffunkytao/SERL-SQL.

Comments18 pages,19 figures, Underreview

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑