LIMIT:少即是多——面向Text-to-SQL的指令微调
LIMIT: Less Is More for Instruction Tuning in Text-to-SQL
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
LIMIT提出以数据为中心的框架,通过难度过滤、思维链合成、质量评分和遗传算法,仅用796和863个样本实现100%表覆盖,使Qwen3-8B在BIRD和Spider上达到69.1%和88.9%执行准确率,超越20倍数据训练的方法,证明数据质量比规模更重要。
AI中文摘要:
大语言模型通过增强推理的微调在Text-to-SQL任务上取得了显著进展,然而现有方法主要依赖大规模指令语料库,并假设数据规模驱动性能提升。我们通过探究一个基本问题来挑战这一范式:有效的Text-to-SQL指令微调所需的最小数据量是多少?我们提出LIMIT(少即是多——面向Text-to-SQL的指令微调),一个以数据为中心的框架,证明了当示例被策略性地选择时,强大的数据库推理能力可以从极其紧凑的训练集中涌现。LIMIT通过四个阶段运作:难度感知过滤,识别模型学习前沿内的样本;思维链合成与基于一致性的选择;通过LLM-as-judge进行多维质量评分;以及联合最大化模式覆盖率和样本质量的遗传算法优化。在BIRD和Spider基准上,LIMIT仅选择796和863个样本,同时实现100%的表覆盖率,使Qwen3-8B达到69.1%和88.9%的执行准确率。这一结果超越了使用20倍数据训练的方法,并在开源方法中确立了新的最先进水平。我们的发现表明,精心的数据策展而非规模,是高效Text-to-SQL学习的关键。
英文摘要:
Large language models have achieved remarkable progress on Text-to-SQL through reasoning-enhanced fine-tuning, yet existing approaches predominantly rely on massive instruction corpora under the assumption that scale drives performance. We challenge this paradigm by investigating a fundamental question: what is the minimal data requirement for effective Text-to-SQL instruction tuning? We propose LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected. LIMIT operates through four stages: difficulty-aware filtering that identifies samples within the model's learning frontier, chain-of-thought synthesis with consistency-based selection, multi-dimensional quality scoring via LLM-as-judge, and genetic algorithm optimization that jointly maximizes schema coverage and sample quality. On the BIRD and Spider benchmark, LIMIT selects only 796 and 863 samples while achieving 100% table coverage, enabling Qwen3-8B to reach 69.1% and 88.9% execution accuracy.This result surpasses methods trained on 20 times more data and establishes a new state-of-the-art among open-source approaches. Our findings suggest that careful data curation, rather than scale, is the key to efficient Text-to-SQL learning.