arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27822cs.DBcs.SE

DBRepro:基于混合约束求解方法的自动化数据库合成框架,用于重现慢查询

DBRepro: Automated Database Synthesis via a Hybrid Constraint-Solving Approach for Reproducing Slow Queries

Zhaoyang Zhang, Shuang Liu, Dengfeng Xu, Wei Lu, Jianquan Leng, Sheng Du, Xiaoyong Du

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对慢查询离线诊断的数据库合成难题,提出DBRepro框架,通过混合约束求解方法生成代理数据库,在TPC-H、SSB及真实数据集上验证了其在基数误差、计划一致性等指标上的优势。

中文摘要 AI 辅助

慢查询常导致数据库管理系统出现严重性能瓶颈。在线诊断其根本原因可能加剧资源争用,而数据隐私法规常禁止将生产数据复制到测试环境。因此,从非侵入式元数据合成代理数据库,使查询优化器生成相同物理执行计划,对离线诊断至关重要。高保真重现需保留全局统计分布,同时强制精确局部基数。现有数据驱动和工作负载感知方法无法同时满足这两个要求。本文提出DBRepro,一种自动化端到端框架,将数据库生成建模为约束分布合成问题。DBRepro从轻量级列统计初始化全局分布,从目标查询提取执行约束,逐步调整分布以满足约束同时保留全局分布。在TPC-H和SSB上的实验表明,与数据驱动基线相比,DBRepro可将基数误差降低多达20.3%,同时保持相同的计划一致性;与工作负载感知基线相比,它可重现多15%的一致执行计划,并将延迟比例误差降低21.5%。我们进一步在人大金仓(KingbaseES)管理的近1TB真实数据集上验证DBRepro,其以高保真度重现复杂慢查询的执行性能。

英文摘要

Slow queries frequently cause severe performance bottlenecks in database management systems. Diagnosing their root causes online risks exacerbating resource contention, while data privacy regulations often prohibit copying production data to test environments. Synthesizing a proxy database from non-intrusive metadata that induces the query optimizer to generate the same physical execution plans is therefore critical for offline diagnosis. High-fidelity reproduction requires preserving global statistical distributions while enforcing exact local cardinalities. Existing data-driven and workload-aware approaches cannot satisfy both requirements simultaneously. We present DBRepro, an automated end-to-end framework that formulates database generation as a constrained distribution synthesis problem. DBRepro initializes a global distribution from lightweight column statistics, extracts execution constraints from target queries, and progressively adjusts the distribution to satisfy these constraints while preserving the global distribution. Experiments on TPC-H and SSB show that DBRepro reduces cardinality error by up to 20.3% over a data-driven baseline while maintaining identical plan consistency. Compared with a workload-aware baseline, it reproduces 15% more consistent execution plans and reduces latency proportion error by 21.5%. We further validate DBRepro on a nearly 1 TB real-world dataset managed by KingbaseES, where it reproduces the execution performance of complex slow queries with high fidelity.

补充信息

↑