AI 中文总结
研究语句级SQL重写,提出SQL-RewriteBench基准测试,应用正确性门控和全分母计算,其指标套件可区分多种指标,定义了SCS和CGOQ,提供多个可执行实例,通过测试发现现有方法存在问题,强调可部署SQL重写需多方面改进。
AI 中文摘要
语句级SQL重写可在不改变数据库管理系统内核的情况下提高查询性能和可维护性,但现有基准测试未将重写方法评估为可部署系统。它们通常关注数据库管理系统性能、规则回归、查询等价性或方言翻译,而忽略了从接受输入查询到生成可执行、结果一致且操作有用的重写的完整路径。我们提出了SQL-RewriteBench,这是一个用于语句级SQL重写的基准测试,它应用了正确性门控和全分母计算。其指标套件明确区分了源接受度、生成率、执行覆盖率、结果一致性、不安全重写率和加速分布。它还定义了SCS(静态SQL结构的确定性索引)和CGOQ(正确性门控优化质量分数),只有在满足特定案例的检查器契约后才给予优化信用。CGOQ通过连续评分函数将运行时改进与结构简化相结合,适用于面向部署的重写评估。作为一个工件,SQL-RewriteBench提供了180个可执行的基准测试实例,分为EQUIV、PERF、ROBUST和DIALECT池,每个实例都打包了SQL、模式元数据、出处、证据和重写机会文档。在七种代表性的学术方法和基于大语言模型的方法中,每个全基准测试的CGOQ都是负的。现有方法往往在重写前失败、结果检查失败或返回正确但比重写输入慢或无更好效果的重写。这些结果表明,可部署的SQL重写需要更广泛的输入处理、结果验证和效益感知的重写决策。
英文摘要
Statement-level SQL rewriting can improve query performance and maintainability without changing the DBMS kernel, but existing benchmarks do not evaluate rewrite methods as deployable systems. They typically focus on DBMS performance, rule regression, query equivalence, or dialect translation, while missing the full path from accepting an input query to producing an executable, result-consistent, and operationally useful rewrite. We present SQL-RewriteBench, a benchmark for statement-level SQL rewriting that applies correctness gating and full-denominator accounting. Its metric suite explicitly separates Source Acceptance, Generation Rate, Execution Coverage, Result Consistency, UnsafeRewrite Rate, and speedup distribution. It also defines SCS, a deterministic index of static SQL structure, and CGOQ, a correctness-gated optimization-quality score that gives optimization credit only after the case-specific Checker Contract is satisfied. CGOQ combines runtime improvement with structural simplification through a continuous scoring function, making it suitable for deployment-oriented rewrite assessment. As an artifact, SQL-RewriteBench provides 180 executable Benchmark Instances organized into EQUIV, PERF, ROBUST, and DIALECT pools, each packaged with SQL, schema metadata, provenance, evidence, and rewrite-opportunity documentation. Across seven representative academic and LLM-based methods, every full-benchmark CGOQ is negative. Existing methods often fail before rewriting, fail result checks, or return correct rewrites that are slower or no better than the input. These results show that deployable SQL rewrite requires broader input handling, result validation, and benefit-aware rewrite decisions.