发表机构
Hainan COSCO Shipping Technology Co., Ltd.(海南中远海运科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对文本到SQL任务,提出CHS-SQL框架,在模式链接阶段结合模型内部置信度进行启发式搜索,平衡精确率和召回率,在SQL生成时进一步优化候选查询,避免局部最优,通过SLM取得了文本到SQL任务的最新成果。
AI 中文摘要
近年来,文本到SQL领域有多项利用小型语言模型(SLM)进行训练的工作。这些方法在生成SQL时仅使用单个NVIDIA RTX 4090 GPU的计算能力,就能达到接近大型模型的性能,同时确保数据安全。多数现有方法在模式链接时过滤冗余表和列以提高文本到SQL的准确性,但未考虑在选择候选模式子集时的精确率-召回率权衡。我们的研究发现模式链接的精确率和召回率直接影响最终SQL准确性。因此,我们提出了CHS-SQL这一新颖框架,在文本到SQL任务上高效微调SLM,不仅平衡精确率和召回率,还提升文本到SQL任务的整体性能。其主要创新在于模式链接阶段,采用结合模型内部置信度的启发式搜索实现最佳精确率-召回率权衡,该精细机制最大化生成SQL查询的相关模式候选的精确率,同时抑制无关噪声。相同策略在SQL生成时进一步应用以优化候选查询,并帮助SLM避免陷入局部最优。我们的方法通过SLM在文本到SQL任务上取得了最新的(SOTA)结果。
英文摘要
Recently, there have been several works in the Text-to-SQL domain that utilize Small Language Models (SLMs) for training. These approaches achieve performance close to that of large models in generating SQL, using only the computational power of a single NVIDIA RTX 4090 GPU, while also ensuring data security. Most existing methods filter out redundant tables and columns during Schema Linking to improve Text-to-SQL accuracy. However, they do not consider the precision-recall trade-off when selecting the candidate schema subset. Our research found that both the precision and recall of Schema Linking directly affect the final SQL accuracy. Therefore, we propose a novel framework for efficiently fine-tuning SLMs on Text-to-SQL tasks, CHS-SQL, that not only balances precision and recall but also improves overall performance on Text-to-SQL tasks. Its main innovation lies in the Schema Linking phase, where a heuristic search combined with model internal confidence is employed to achieve an optimal precision-recall trade-off. This elaborated mechanism maximizes the precision of relevant schema candidates for the generated SQL queries while suppressing irrelevant noise. The same strategy is further applied during SQL generation to refine candidate queries while helping the SLM to avoid trapping in a local optimum. Our method achieves state-of-the-art (SOTA) results on Text-to-SQL tasks via SLMs.