AI 中文总结
研究SQL模式检索问题,提出语料库自适应微调方法,通过合成查询、挖掘负样本及对比微调嵌入器,提升了检索性能,使模型成为十亿参数下最强检索器,确立模式链接为独立任务及语料库适应为实用部署途径。
AI 中文摘要
在SQL环境中的检索主要被研究为在大量SQL语句中找到回答自然语言问题的语句的任务。然而,在大规模情况下,一个更基本的检索问题先于生成:模式检索,即在可能包含数千个表和列的数据库中识别问题所需的表和列,这远远超出了模型上下文的容纳范围。我们认为这一步骤值得进行一流的评估。为此,我们将五个文本到SQL数据集(Spider、BIRD、BEAVER和两个LiveSQLBench变体)重新构建为表和列粒度的检索任务,涵盖两种文档表示下的现实和企业规模模式,并且我们表明现成的文本和代码嵌入器在这种设置下转移效果不佳。然后我们提出语料库自适应微调:直接从目标模式语料库中合成自然语言查询,挖掘粒度感知的硬负样本,并对一个3亿500万参数的嵌入器进行对比微调。此过程将平均召回率@10从60.4提高到75.6(归一化折损累计增益@10从51.9提高到68.0),使3亿500万参数的模型成为十亿参数以下最强的检索器,并且与40亿到80亿参数的最先进嵌入器具有竞争力,后者比它大一个数量级以上。相同的方法将一个80亿参数的最先进嵌入器的召回率@10从77.8提高到78.4,与基准测试中的最佳结果相匹配,表明这种适应与主干无关。留一语料库实验和泄漏审计表明,这些收益反映了可转移的模式检索能力,而不是对评估数据的记忆。我们的结果将模式链接确立为一个独立的检索任务,并将轻量级、无标签的语料库适应确立为在企业规模上部署它的实用途径。
英文摘要
Retrieval in the SQL setting has largely been studied as the task of finding, within a large collection of SQL statements, the statement that answers a natural-language question. At scale, however, a more fundamental retrieval problem precedes generation: schema retrieval, identifying the tables and columns a question requires in a database that may contain thousands of them, far more than fit in a model's context. We argue that this step warrants first-class evaluation. To this end, we recast five text-to-SQL datasets (Spider, BIRD, BEAVER, and two LiveSQLBench variants) as retrieval tasks at both table and column granularity, covering realistic and enterprise-scale schemas under two document representations, and we show that off-the-shelf text and code embedders transfer poorly to this setting. We then propose corpus-adaptive fine-tuning: natural-language queries are synthesized directly from the target schema corpus, granularity-aware hard negatives are mined, and a 305M-parameter embedder is fine-tuned contrastively. This procedure raises average recall@10 from 60.4 to 75.6 (nDCG@10 from 51.9 to 68.0), making the 305M model the strongest retriever under one billion parameters and competitive with state-of-the-art embedders of 4-8B parameters, more than an order of magnitude larger. The same recipe improves an 8B state-of-the-art embedder from 77.8 to 78.4 recall@10, matching the best result on the benchmark and indicating that the adaptation is backbone-agnostic. Leave-one-corpus-out experiments and a leakage audit show that these gains reflect a transferable schema-retrieval ability rather than memorization of the evaluation data. Our results establish schema linking as a standalone retrieval task and lightweight, label-free corpus adaptation as a practical route to deploying it at enterprise scale.