算术可变对数对数:推进内存-方差前沿
Arithmetic Variable LogLog: Advancing the Memory-Variance Frontier
浏览论文内容
中文总结 AI 辅助
针对数据流基数估计的内存-方差权衡问题,提出AVLL方法,通过算术编码等技术超越ExaLogLog,精度提升4.7%,速度快2.7-4.5倍,内存规模更优且适配多种数据分布。
中文摘要 AI 辅助
基数估计(即统计数据流中不同元素的数量)需要在内存与精度之间进行权衡。ExaLogLog 近期通过将宽寄存器与费雪信息最优的最大似然(ML)估计器相结合,在这种权衡中达到了当前最优水平,在 HyperLogLog 变体中实现了已知最优的内存-方差乘积(MVP)。本文提出算术可变对数对数(Arithmetic Variable LogLog,AVLL),该方法通过算术编码并消除不常见状态,在所有测试的内存点上均超越 ExaLogLog,每 11 个寄存器即可完全利用 64 位字,带来 5.5 倍的寄存器数量优势。其四分量混合估计器 HLDLC 利用这种密度优势,无需迭代求解即可超越 ExaLogLog 的 ML 估计精度。在 1 KB 时,AVLL 的宽度加权平均绝对误差为 1.63%,而 ExaLogLog 为 1.71%,提升了 4.7%;对应的经验 MVP 为 3.4,超越 ExaLogLog 的实际 MVP(3.78)及其理论最优值(3.67),且在 0.25 KB 至 4 KB 的所有测试规模下均成立。AVLL 继承了 DynamicLogLog 的早退出机制,该机制在触及任何寄存器前过滤掉大部分元素;在每个线程同时运行数千个草图时,由于早退出减少了内存带宽需求,AVLL 比 ExaLogLog 快 2.7 至 4.5 倍。与 DynamicLogLog 类似,AVLL 采用共享偏移存储相对 NLZ 值,因此其内存规模为 O(B + log log C),而非 O(B × log log C),从而将最大可表示基数与寄存器宽度解耦。这些结果在高复杂度(全唯一)和低复杂度(非均匀高重复率)数据分布下均成立,且重复不会导致精度下降。AVLL 被实现为一个独立的 Java 类,内置所有校正公式,可在 BBTools 套件中获取,链接为 this https URL。
英文摘要
Cardinality estimation - counting the number of distinct elements in a data stream - requires a tradeoff between memory and accuracy. ExaLogLog recently established the state of the art for this tradeoff by combining wide registers with a Fisher-information-optimal maximum likelihood (ML) estimator, achieving the best known memory-variance product (MVP) among HyperLogLog variants. Here we present Arithmetic Variable LogLog (AVLL), which surpasses ExaLogLog at every memory point tested using arithmetic encoding and eliminating uncommon states to consume 64-bit words completely with 11 registers each, yielding a 5.5x register-count advantage. Its four-component blended estimator, HLDLC, exploits this density advantage to surpass ExaLogLog's ML accuracy without iterative solving. At 1 KB, AVLL achieves 1.63% width-weighted mean absolute error compared to ExaLogLog's 1.71% - a 4.7% improvement. The corresponding empirical MVP is 3.4, surpassing ExaLogLog's practical MVP of 3.78 and its theoretical optimum of 3.67. This holds at every tested size from 0.25 to 4 KB. AVLL inherits DynamicLogLog's early exit mechanism, which filters most elements before any register is touched. With thousands of simultaneous sketches per thread, AVLL is 2.7-4.5x faster than ExaLogLog due to the reduced memory bandwidth from early exits. Like DynamicLogLog, AVLL stores relative NLZ values with a shared offset, so its memory scales as O(B + log log C) rather than O(B x log log C) - decoupling maximum representable cardinality from register width. These results hold under both high-complexity (all-unique) and low-complexity (nonuniformly high duplication rate) data distributions, with zero accuracy degradation from duplication. AVLL is implemented as a single self-contained Java class with all correction formulas embedded, available in the BBTools suite at https://bbmap.org.