arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ModularSQL:文本到SQL中多重性盲点的运行时防护栏

ModularSQL: A Runtime Guardrail for the Multiplicity Blind Spot in Text-to-SQL

Tianxin Zhou, Ruixi Lin

arXiv 2609.29573首次发表:更新:

发表机构

University of Southern California; Northeastern University(南加州大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本到SQL系统在基准评估中忽略多重性错误的问题,提出ModularSQL运行时防护栏,通过Multiset-EX评估暴露盲点,并利用确定性补丁和低成本LLM救援,以微小开销提升执行安全性。

AI 中文摘要

文本到SQL系统越来越多地部署在生产数据库上,其中通过基准评估的查询仍可能产生扭曲下游工作流的结果。标准的基于集合的执行准确率(Set-EX)会折叠重复行,因此可能遗漏多重性错误,包括缺失DISTINCT、膨胀的聚合以及笛卡尔式连接爆炸。我们将此称为多重性盲点(MBS),并引入Multiset-EX,一种保留多重性的评估标准,以暴露此类失败。在来自三个骨干模型(Qwen2.5-Coder-32B、Qwen3-Coder-30B-A3B和Gemma-3-27B)的已发布DeepEye-SQL工件上,在可执行的BIRD-Dev N=1532上,我们发现Set-EX和Multiset-EX之间存在一致的5.81--6.79个百分点差距。该差距并非DeepEye-SQL特有:它在已发布的DAIL-SQL+GPT-4(5.22个百分点)和BIRD GPT-3.5-turbo(3.39个百分点)预测上持续存在。我们进一步引入ModularSQL,一种轻量级的后选择运行时防护栏,它探测执行结果中的多重性异常,并仅对标记的查询应用确定性补丁或低成本LLM救援。集成DeepEye-SQL使用Qwen3-Coder,ModularSQL将Set-EX保持在72.06%,同时将Multiset-EX从65.86%提高到67.75%(+1.89个百分点)。它标记了77个高风险异常,同时仅增加0.0076美元的总LLM成本和每次查询120毫秒的摊销延迟。跨管道评估表明,无候选检测器和确定性补丁也迁移到独立发布的预测集。总体而言,这些结果表明基准准确率并不一定意味着执行安全的SQL,而轻量级、多重性感知的运行时防护栏可以以适度的计算开销缩小这一差距。

英文摘要

Text-to-SQL systems are increasingly deployed on production databases, where queries that pass benchmark evaluation can still produce results that distort downstream workflows. Standard set-based execution accuracy (Set-EX) collapses duplicate rows and can therefore miss multiplicity errors, including missing DISTINCT, inflated aggregates, and Cartesian-style join explosions. We call this the Multiplicity Blind Spot (MBS) and introduce Multiset-EX, a multiplicity-preserving evaluation criterion that exposes such failures. Across released DeepEye-SQL artifacts from three backbones (Qwen2.5-Coder-32B, Qwen3-Coder-30B-A3B, and Gemma-3-27B) on executable BIRD-Dev N=1532, we find a consistent 5.81--6.79 pp gap between Set-EX and Multiset-EX. The gap is not specific to DeepEye-SQL: it persists on released DAIL-SQL+GPT-4 (5.22 pp) and BIRD GPT-3.5-turbo (3.39 pp) predictions. We further introduce ModularSQL, a lightweight post-selection runtime guardrail that probes executed results for multiplicity anomalies and applies deterministic patches or low-cost LLM rescue only to flagged queries. Integrated with DeepEye-SQL using Qwen3-Coder, ModularSQL preserves Set-EX at 72.06% while improving Multiset-EX from 65.86% to 67.75% (+1.89 pp). It flags 77 high-risk anomalies, while adding only $0.0076 in total LLM cost and 120 ms amortized latency per query. Cross-pipeline evaluation shows that the candidate-free detector and deterministic patches also transfer to independently released prediction sets. Overall, these results show that benchmark accuracy does not necessarily imply execution-safe SQL, and that lightweight, multiplicity-aware runtime guardrails can narrow this gap with modest computational overhead.

Comments12 pages, 5 tables, 4 figures. Code: https://github.com/Ruixi1313/ModularSQL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑