arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用记忆替代训练:面向Text-to-SQL的列表式选择方法

Replacing Training with Memory: Listwise Selection for Text-to-SQL

Yeonseok Jeong, Soyoung Yoon, Seongjun Lee, Seung-won Hwang

arXiv 2609.00834首次发表:更新:

发表机构

Seoul National University; KAIST; IPAI(首尔大学; 韩国科学技术院; IPA研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无需微调的列表式选择器MaP-SQL,以记忆替代微调目标,在Text-to-SQL基准上较R^3-SQL提升执行准确率且减少令牌数,实现更稳定高效的选择。

AI 中文摘要

现代Text-to-SQL系统常采用生成-执行-选择的流水线,生成多个候选查询后选择最优者。列表式选择通过联合比较多个候选已被广泛采用,但微调列表式选择器的成本很高。因此我们提出一种无需微调的列表式选择器,用推理时策略替代两大微调目标:(1)将选择准则学习为排序;(2)缓解位置偏差。首先,我们构建可复用的结构化记忆,而非将选择行为学习为模型参数。给定问题,MaP-SQL会检索从训练数据蒸馏得到的记忆,这些记忆编码了自然语言如何映射到模式元素、SQL操作及预期输出,作为以列表式方式评估候选的显式决策准则。其次,为缓解列表式选择器的排序偏差,我们聚合多个输入排列的排序结果,通过执行结果和点式评分优化推理成本。我们的方法在保持效率及与现有大语言模型兼容性的同时提升了选择准确率。在Text-to-SQL基准测试中,它无需微调即可产生更稳定的选择,且比现有方法的不必要比较更少。在BIRD-dev上,使用相同候选集时,它平均比此前基于选择器的最优方法R^3-SQL的执行准确率高出2.02个百分点,且令牌数减少2.92倍。

英文摘要

Modern Text-to-SQL systems often follow generate-execute-select pipelines, generating multiple candidate queries then selecting the best one. Listwise selection, by jointly comparing multiple candidates, has been widely adopted, but fine-tuning listwise selectors is costly. We thus propose a fine-tuning-free listwise selector. We replace two major fine-tuning objectives with inference-time strategies: (1) learning selection criteria as ordering and (2) mitigating positional bias. First, we build reusable structured memories instead of learning selection behavior as model parameters. Given a question, MaP-SQL retrieves memories distilled from training data that encode how natural language maps to schema elements, SQL operations, and expected outputs. These memories serve as explicit decision criteria for evaluating candidates in a listwise manner. Second, to mitigate ordering bias of listwise selectors, we aggregate rankings across multiple input permutations, with inference cost optimized by execution results and pointwise scoring. Our approach improves selection accuracy while maintaining efficiency and compatibility with existing large language models. Across Text-to-SQL benchmarks, it produces more stable selection without fine-tuning and fewer unnecessary comparisons than existing methods. On BIRD-dev, it outperforms the previous state-of-the-art selector-based method R^3-SQL by 2.02 execution accuracy points on average using the same candidate sets, with 2.92x fewer tokens.

CommentsAccepted by Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑