AI 中文总结
研究人员提出Diva++范围过滤器,解决现有范围过滤器的三类缺陷,通过优化中缀存储实现最优内存与误报率权衡,支持动态更新及变长键与查询,性能达到现有最优水平。
AI 中文摘要
范围过滤器是一种紧凑的概率数据结构,用于回答近似范围空值查询,被应用于诸多领域,例如键值存储系统中,可快速排除给定查询范围内键的存在性,避免在存储中对其进行搜索。然而,现有的所有范围过滤器均存在至少一项缺陷:(1)不提供任何误报率或性能保证;(2)不支持变长键和查询范围;(3)不支持动态更新。我们提出Diva,这是首个同时解决上述所有挑战的范围过滤器。Diva通过采样键并将其存储在缓存高效的字典树(trie)中,学习数据集的分布;它通过去除样本间键的最长公共前缀、截断后缀,同时保留中间足够的位(即中缀)以按排序顺序区分键,来压缩键;它将中缀存储在常数时间动态数据块中,通过拆分数据块处理插入和扩展操作;它通过遍历字典树并检查中缀是否包含在目标查询范围内来处理范围查询。我们从数学上证明,在许多常见的真实世界数据分布上,Diva在内存和误报率之间实现了最优权衡。我们通过引入增强型Diva变体Diva++,将这些优势扩展到更广泛的真实工作负载中。Diva++通过使用保序熵编码去除中缀间的冗余来节省内存,随后移除所有剩余的相同中缀,并利用释放的空间在紧凑二进制字典树中存储更多原始键的位。我们将Diva和Diva++与所有现有范围过滤器进行比较,结果显示,在真实世界数据集上,它们的误报率与现有最优水平相当,同时支持动态性以及变长查询和键。
英文摘要
Range filters are compact probabilistic data structures that answer approximate range emptiness queries. They are used in many domains, e.g., in key-value stores, to quickly rule out the existence of keys in a given query range and avoid searching for them in storage. However, all existing range filters exhibit at least one of three shortcomings: (1) they do not provide any false positive rate or performance guarantees, (2) they do not support variable-length keys and query ranges, and (3) they do not allow dynamic updates. We introduce Diva, the first range filter to address all the above challenges simultaneously. Diva learns the dataset's distribution by sampling keys and storing them in a cache-efficient trie. It compresses keys in-between samples by removing their longest common prefix and truncating their suffixes while leaving enough bits in the middle (i.e., an infix) to differentiate the keys in sorted order. It stores infixes in constant-time dynamic data blocks, which it splits to handle insertions and expansions. It processes a range query by traversing the trie and checking for the inclusion of infixes in the target query range. We mathematically prove that Diva provides the best possible trade-off between memory and false positive rate on many common real-world data distributions. We extend these benefits to a wider range of real-world workloads by introducing Diva++, an enhanced Diva variant. Diva++ saves memory by removing redundancies among infixes using order-preserving entropy encoding. It then removes any remaining identical infixes and uses the freed space to store more bits of the original keys within compact binary tries. We compare Diva and Diva++ to all prior range filters, and show that they achieve a false positive rate on par with the state of the art on real-world datasets while supporting dynamicity and variable-length queries and keys.
Comments31 pages. 19 figures. 6 tables. Submitted to the PVLDB Journal