arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21133cs.DBcs.AI

随机转变:面向AI操作符的Text-to-SQL新评估范式

The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators

发表机构谷歌
查看机构详情
  • Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

Tarfah Alrashed, Fatma Ozcan, Per Jacobsson, Tal Neiman, Xianshun Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种多层评估框架,将确定性数据库逻辑与灵活AI语义解耦,分别验证关系逻辑与AI操作,以解决Text-to-SQL中AI操作符非确定性输出的评估难题,在BigQuery和ThalamusDB上实现高达97.2%的整体准确率。

中文摘要 AI 辅助

SQL已通过AI操作符得到增强,使现代数据分析平台能够从结构化和非结构化数据中获取洞察。我们观察到,尽管当前的Text-to-SQL系统能够成功生成这些AI增强查询,但可靠地评估其正确性仍然是一个关键且开放的挑战。当前依赖于精确查询结果和确定性执行的评估指标,在面对AI操作符灵活且非确定性的输出时系统性地失效。在本文中,我们形式化了这些独特的评估失败模式,并引入了一个多层评估框架,该框架将确定性数据库逻辑与灵活的AI语义解耦。我们在工业界(BigQuery)和学术界(ThalamusDB)的系统中测试了我们的方法。我们证明,传统的执行准确率严重惩罚有效查询,对正确翻译的检测率低至25%。此外,即使是最先进的基于LLM的自动评审器,由于同时判断关系组件和AI组件的复杂性,也会错误地拒绝32%的准确查询。通过分别验证标准关系逻辑和AI操作,我们的框架在两个平台上均达到了最先进的整体准确率(高达97.2%),为基准测试AI驱动的SQL生成器提出了一个可靠的基准。

英文摘要

SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights from both structured and unstructured data. We observe that while current Text-to-SQL systems can successfully generate these AI-augmented queries, reliably evaluating their correctness remains a critical open challenge. Current metrics, which rely on exact query results and deterministic execution, systematically fail against the flexible, non-deterministic outputs of AI operators. In this paper, we formalize these unique evaluation failure modes and introduce a Multilayered Evaluation Framework that decouples deterministic database logic from flexible AI semantics. We test our approach across both industry (BigQuery) and academic (ThalamusDB) systems. We demonstrate that traditional Execution Accuracy severely penalizes valid queries, achieving as low as a 25% detection rate for correct translations. Furthermore, even a state-of-the-art LLM-based autorater falsely rejects 32% of accurate queries due to the complexity of judging both relational and AI components simultaneously. By validating the standard relational logic and the AI operations separately, our framework achieves state-of-the-art overall accuracy across both platforms (up to 97.2%), proposing a reliable standard for benchmarking AI-powered SQL generators.

↑