arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18135cs.CLcs.AIcs.DB

DualSQL: 基于多智能体强化学习的文本到SQL生成

DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning

Shijie Chen, Yu Gan, Yeounoh Chung, Jiani Zhang, Quannan Li, Sravan Babu Bodapati, Cody J. Greer, Yu Su, Fatma Ozcan

首次发表
浏览论文内容

中文总结 AI 辅助

DualSQL提出一种基于共享模型骨干的双智能体强化学习框架,通过联合优化模式链接与SQL生成,并引入REX度量,以少量数据训练达到超越32B模型的性能。

中文摘要 AI 辅助

最先进的文本到SQL系统通常是围绕两个基本任务构建的多智能体流水线:模式链接和SQL生成。然而,现有工作为每个任务训练单独的模型,未能利用这些相互关联任务之间的协同效应。在本工作中,我们提出了DualSQL,一种新的文本到SQL系统,由两个智能体组成,它们共享同一个模型骨干。这两个智能体共享相同的模型权重和智能体框架,通过一个稳健的多智能体强化学习(RL)框架实现联合优化。我们设计了三个数据库访问工具,以促进基于数据库交互的多步推理。为了改进训练并避免模型崩溃,我们引入了一组rollout护栏机制,稳定多智能体RL训练,支持DualSQL在训练过程中持续改进。我们还引入了一个新的SQL正确性度量——稳健执行匹配(REX),以更准确地判断SQL正确性并分配奖励信号。仅使用3755个示例进行训练,DualSQL-4B在BIRD开发集上达到了68.0%的执行准确率,与之前的7B模型相当。DualSQL-8B进一步提高到71.1%,超过了之前具有32B参数的最先进单模型解决方案。这些结果证明了联合多智能体强化学习在构建高性能文本到SQL流水线方面的优势。

英文摘要

State-of-the-art Text-to-SQL systems are typically multi-agent pipelines centered around two fundamental tasks: schema linking and SQL generation. However, existing work trains separate models for each task, failing to leverage the synergy between these interrelated tasks. In this work, we propose DualSQL, a new Text-to-SQL system consisting of two agents powered by a single model backbone. The agents share the same model weights and agentic scaffold, enabling joint optimization through a robust multi-agent reinforcement learning (RL) framework. We design three database access tools to facilitate effective multi-step reasoning grounded to interactions with the databases. To improve training and avoid model collapse, we introduce a set of rollout guardrail mechanisms that stabilizes multi-agent RL training, supporting DualSQL to keep improving during training. We also introduce a new SQL correctness metric, robust execution match (REX), to more accurately judge SQL correctness and assign reward signals. Being trained on only 3755 examples, DualSQL-4B achieves an impressive 68.0% execution accuracy on the BIRD development set, matching previous 7B models. DualSQL-8B further improves to 71.1%, outperforming previous state-of-the-art single-model solutions with 32B parameters. These results demonstrate the strength of joint multi-agent reinforcement learning for building high performance Text-to-SQL pipelines.

发表机构

  • The Ohio State University(俄亥俄州立大学)
  • Google LLC(谷歌有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

↑