AI 中文总结
研究满足函数依赖的数据库上连接查询答案字典序直接访问的复杂性,给出预处理时间上下界。先介绍简单方法及不足,再阐述受信息论启发的方法,虽界不紧,但刻画了线性预处理时间的字典序直接访问,下界依赖零团猜想且仅适用于无自连接查询。
AI 中文摘要
我们研究了在满足函数依赖(FD)的数据库上对连接查询答案进行字典序直接访问的复杂性。具体而言,我们给出了实现多对数访问时间所需预处理时间的细粒度上下界。首先考虑重新排序扩展的简单方法,它先将FD纳入查询和排序,然后在评估时忽略FD,结果表明该方法对一元FD给出了紧界,但对一般FD失败。接着考虑第二种方法,受利用信息论对查询答案大小界的启发,在为手头直接访问任务定制的分解的包实例化时考虑FD,有趣的是,在构建分解时相同的重新排序也有助于降低复杂性。虽然得到的上下界通常不紧,但它们给出了具有线性预处理时间的字典序直接访问的完整刻画。本文所有下界仅适用于无自连接的查询,并依赖于零团猜想。
英文摘要
We study the complexity of lexicographic direct access to join query answers over databases that satisfy functional dependencies (FDs). More precisely, we give fine-grained lower and upper bounds on the preprocessing time required to achieve polylogarithmic access time. We start by considering the simple approach of a reordered extension, which first incorporates the FDs in the query and order, and then ignores the FDs during evaluation. We show that this simple approach gives tight bounds for unary FDs but fails for general FDs. We then consider a second approach, inspired by size bounds for query answers using information theory, that takes the FDs into account while materializing the bags of a decomposition tailored to the direct access task at hand. Interestingly, we show that the same reordering is also useful while constructing the decomposition in this second approach for reducing the complexity. While the obtained upper and lower bounds are generally not tight, we show that they yield a complete characterization of lexicographic direct access with linear preprocessing time. All lower bounds in this paper apply only to queries without self-joins and rely on the Zero-Clique Conjecture.