QEmbed:一种基于深度学习的高效查询处理基数估计器
QEmbed: A Deep Learning Based Cardinality Estimator for Efficient Query Processing
- Indian Institute of Technology Jammu(詹穆印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有基数估计方法在混合高/低基数属性数据集上难以平衡内存与精度的问题,提出基于MADE自回归框架的QEmbed深度学习模型,采用混合编码捕获属性相关性,有效降低极端尾部误差。
AI中文摘要:
基数估计是任何商业数据库系统高效查询处理的核心。几十年来,基于非学习的估计技术(如基于直方图、基于采样的方法)已广泛用于商业和开源数据库平台。然而,这些技术仅在表中列数较少时有效,因为它们无法正确捕获多个属性之间的依赖关系。近年来,基于学习的方法已被证明显著优于过去三十年使用的启发式方法。尽管取得了这些成功,现有的学习模型在处理混合高基数和低基数属性的数据集时,往往难以平衡内存效率和准确性。在本文中,我们提出了一种正式称为QEmbed的深度学习模型。我们的模型建立在用于分布估计的掩码自编码器(MADE)自回归框架之上,以学习用于选择性估计的联合数据分布。为了改进数据表示并克服使用单一编码方法的局限性,我们设计了一种结合独热编码和嵌入编码的混合编码方案。这种混合设计使QEmbed能够为较小的域保留细粒度的属性信息,同时为大型稀疏域捕获紧凑的语义模式。我们通过将联合数据分布分解为一系列条件概率来捕获属性相关性。这种方法自然适用于点查询和范围查询。通过大量实验,我们表明,虽然QEmbed在极宽模式上面临延迟权衡,但总体上它提供了高度可靠的基数估计。我们模型的一个关键优势是它减少了极端尾部误差(最大Q误差),避免了在复杂、高度相关的工作负载上发生灾难性的估计失败。
英文摘要:
Cardinality estimation is at the core of any commercial database system for efficient query processing. Over the decades, non-learning-based estimation techniques (e.g., histogram-based, sampling-based) have been widely used in both commercial and open-source database platforms. However, these techniques are only effective when the number of columns in a table is small, as they cannot properly capture dependencies between multiple attributes. Recently, learning-based approaches have been shown to perform significantly better than the heuristic methods that have been used for the past three decades. Despite this success, existing learned models often struggle to balance memory efficiency and accuracy when dealing with datasets that mix high and low cardinality attributes. In this paper, we propose a deep learning model formally called QEmbed. Our model is built upon the Masked Autoencoder for Distribution Estimation (MADE) auto-regressive framework to learn joint data distributions for selectivity estimation. To improve data representation and overcome the limitations of using a single encoding method, we design a hybrid encoding scheme that combines one-hot and embedding encodings. This hybrid design enables QEmbed to retain fine-grained attribute information for smaller domains while capturing compact semantic patterns for large, sparse domains. We capture attribute correlations by factoring the joint data distribution into a series of conditional probabilities. This approach naturally accommodates both point and range queries. Through extensive experiments, we show that while QEmbed faces a latency trade-off on extremely wide schemas, it provides highly reliable cardinality estimates overall. A key advantage of our model is that it reduces extreme tail errors (maximum Q-errors), avoiding catastrophic estimation failures on complex, highly correlated workloads.