arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FALCON:面向NL2SQL对的合成数据生成的模型与数据集无关框架

FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs

Darian Lee, Shannon Rumsey, Jack St. Clair, Xinyi Tang, Aditya Bansal, Yuanming Shi

arXiv 2610.03625首次发表:更新:

发表机构

University of California, Santa Cruz; Adobe(加州大学圣克鲁兹分校; Adobe公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FALCON提出一种模型与数据集无关的框架,利用保留字SQL种子和角色提示生成复杂度高、歧义感知的合成NL2SQL数据,并通过对齐过滤保留难度,实验证明其训练模型在复杂查询上优于基线。

AI 中文摘要

关系数据库是结构化知识最广泛部署的形式之一,而对其的自然语言访问需要将语言映射到模式实体和关系上,同时处理人们在表达请求时固有的歧义。现有的合成NL-to-SQL数据生成方法大多忽略这种歧义,产生过于简化的查询,无法让模型为现实世界结构化知识访问的复杂性做好准备。我们提出FALCON,一个能够生成与具有挑战性的现实基准复杂度相匹配的、逼真的、具有歧义意识的NL-to-SQL数据的框架,且使用紧凑的开源模型以低成本实现。我们的方法结合保留字SQL种子和基于角色的提示来生成结构复杂的查询,而基于对齐的过滤通过区分真正错误的示例与复杂但有效的查询来保留难度。人工评估确认了跨模型规模的持续高质量,我们生成的数据在SQL复杂度和自然语言丰富度上均超过现有基准。难度分层分析表明,在FALCON数据上训练的模型随着查询复杂度的增加,其性能越来越优于基线训练的模型,验证了我们的流程在生成具有挑战性的训练数据方面的成功。当与一小部分现有基准数据结合时,混合训练恢复了在较简单查询上的性能,同时在复杂查询上保留了这些优势。模型和数据库无关的设计使组织能够在本地生成高复杂度的NL-to-SQL训练数据,而无需外部API。

英文摘要

Relational databases are among the most widely deployed forms of structured knowledge, and natural language access to them requires grounding language onto schema entities and relations while handling the ambiguity inherent in how people phrase requests. Existing synthetic NL-to-SQL data generation methods largely ignore this ambiguity and produce oversimplified queries that fail to prepare models for the complexity of real-world structured knowledge access. We present FALCON, a framework that generates realistic, ambiguity-aware NL-to-SQL data matching the complexity of challenging real-world benchmarks, at low cost using compact open models. Our approach combines reserved-word SQL seeding and persona-based prompting to generate structurally complex queries, while alignment-based filtering preserves difficulty by distinguishing genuinely incorrect examples from complex but valid queries. Human evaluation confirms consistent high quality across model sizes, and our generated data exceeds existing benchmarks in both SQL complexity and natural language richness. Difficulty-stratified analysis shows models trained on FALCON data increasingly outperform baseline-trained models as query complexity increases, validating our pipeline's success in generating challenging training data. When combined with a small proportion of existing benchmark data, mixed training recovers performance on simpler queries while preserving these advantages on complex ones. The model- and database-agnostic design enables organizations to generate high-complexity NL-to-SQL training data locally without external APIs.

CommentsAccepted to AKBC Workshop, EMNLP

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑