arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于物化视图的查询重写全流程基准测试

Benchmarking the Full Pipeline of Materialized-View-Based Query Rewriting

Xinjie Hu, Zhengjie Miao

arXiv 2607.19679首次发表:更新:

AI 中文总结

本文用模块化框架和可控消融基准测试基于物化视图的查询重写全流程,引入跨引擎协议对比不同系统。发现跨阶段交互效应及物化视图使用差异,识别重写后性能回归模式,突出限制性能阶段,为相关设计提供依据。

AI 中文摘要

物化视图通过预计算可重用子表达式来加速OLAP和数据仓库工作负载,但基于物化视图的实际查询加速是一个多阶段流程:候选枚举、存储预算下的视图选择以及优化器内的查询重写。现有评估通常仅研究此流程的部分内容且局限于单个系统。本文通过模块化评估框架联合评估枚举、选择和重写,并使用可控消融来对基于物化视图的查询重写进行基准测试。还引入了跨引擎协议,通过对比原生优化器级重写和可用的便携式SQL重写基线来比较仅公开执行计划的系统。研究发现跨阶段存在强交互效应,物化视图使用和实现的节省存在很大差异,识别出导致重写后性能回归的常见失败模式。结果突出了最常限制性能的流程阶段,为指导未来物化视图枚举、选择和重写设计提供了依据。

英文摘要

Materialized views (MVs) accelerate OLAP and data-warehouse workloads by precomputing reusable subexpressions, but practical MV-based query acceleration is a multi-stage pipeline: candidate enumeration, view selection under storage budgets, and query rewriting inside the optimizer. Existing evaluations typically study only parts of this pipeline and within a single system, leaving end-to-end trade-offs and cross-system behavior unclear. In this paper, we benchmark MV-based query rewriting by jointly evaluating enumeration, selection, and rewriting with a modular evaluation framework and by using controlled ablations. We also introduce a cross-engine protocol allowing us to compare systems that expose only execution plans by contrasting native optimizer-level rewriting with portable SQL rewriting baselines when available. Across representative academic methods and modern open-source and commercial systems, we find strong interaction effects across stages and large variability in MV usage and realized savings. We identify recurring failure modes that explain performance regressions after rewriting. Our results highlight which pipeline stages most often limit performance and provide evidence to guide future MV enumeration, selection, and rewriting designs.

Journal refProc. VLDB Endow. 19, 11 (2026), 3772-3785

DOI:10.14778/3836663.3836724

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑