arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弥合跨方言差距:查询计划作为文本到SQL中的可移植接口

Closing the Cross-Dialect Gap: Query Plans as a Portable Interface in Text-to-SQL

Corentin Royer, Robin Oester, Yotam Perlitz, Yannick Metz, Andrea Giovannini, Mennatallah El-Assady

arXiv 2609.33670首次发表:更新:

发表机构

IBM Research; ETH Zurich(IBM研究; 苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本到SQL跨方言性能下降问题,提出以方言无关的关系代数查询计划为生成目标,经编译器转换,在十三个模型上恢复可移植性,并引入MetricName公平评估,证明计划监督优于SQL监督。

AI 中文摘要

文本到SQL系统通常在单一方言(SQLite)上训练和评估,但生产部署却涵盖PostgreSQL、MySQL、ClickHouse等更多数据库。我们表明,这种单一方言假设导致我们测试的每个模型在跨方言准确性上出现显著下降。这种下降在不同规模、架构甚至专门构建的文本到SQL系统中均持续存在。我们认为,解决方案在于改变生成目标:不再要求大型语言模型(LLM)输出特定方言的SQL,而是让其生成与方言无关的关系代数查询计划,再由确定性编译器将其转换为任何受支持后端的SQL。在从3B到前沿规模的十三个模型上,这种方法几乎均匀地恢复了跨方言可移植性,对于有能力的提示模型,在其母语方言上的峰值准确性仅有小幅损失,而一旦在计划上进行微调则无损失;在匹配的微调条件下,计划监督产生的模型优于SQL监督。我们还引入了MetricName,一个问题感知的结果集比较器,用于跨方言公平评估,因为现有指标将语义错误与良性的跨方言变异混为一谈。更广泛地说,这一结果提醒我们,为执行而选择的生成目标不一定能最大化生成质量。

英文摘要

Text-to-SQL systems are typically trained and evaluated on a single dialect (SQLite), yet production deployments span PostgreSQL, MySQL, ClickHouse, and beyond. We show that this single-dialect assumption leads to a substantial drop in cross-dialect accuracy for every model we tested. The drop persists across scale, architecture, and even purpose-built text-to-SQL systems. We argue that the fix is to change the generation target: instead of asking an LLM to emit dialect-specific SQL, we have it emit a dialect-agnostic relational algebra query plan, which a deterministic compiler then renders into SQL for any supported backend. Across thirteen models from 3B to frontier scale, this restores cross-dialect portability nearly uniformly, at a small cost in peak accuracy on the model's home dialect for capable prompted models and none once fine-tuned on plans; under matched fine-tuning, plan supervision yields a stronger model than SQL supervision. We also introduce MetricName, a question-aware result-set comparator needed to evaluate fairly across dialects, where existing metrics confound semantic errors with benign cross-dialect variation. More broadly, the result is a reminder that a generation target chosen for execution is not necessarily the one that maximizes generation quality.

CommentsAccepted at Findings of the Association for Computational Linguistics: EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑