发表机构
Université Paris Nanterre(巴黎楠泰尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从概率视角重新分析经典概率算法HyperLogLog,通过基本方法建立指数偏差不等式,实现对海量数据集中不同元素数量的估计,其估计非渐近且完全明确。
AI 中文摘要
HyperLogLog是一种经典概率算法,用于估计海量数据集中不同元素的数量,只需遍历一次数据。原始论文对其输出的期望和方差进行了精确分析。本文从概率视角重新审视该算法,建立了指数偏差不等式。方法简单,估计非渐近且完全明确。
英文摘要
HyperLogLog is a now classic probabilistic algorithm that provides an approximation of the number of distinct elements in a massive dataset, using only one pass over the data. In the original article, Flajolet, Fusy, Gandouet, Meunier (2007) provided a sharp analysis of the expectation and variance of the output, using explicit formulas analyzed using poissonization and Mellin transform. In this short article, we revisit the analysis of HyperLogLog with a more probabilistic viewpoint. This allows us to establish exponential deviation inequalities for the HyperLogLog estimator. The methods are elementary, but the estimates are non-asymptotic and totally explicit.
CommentsUpdated bibligraphy