AI 中文总结
该研究证明了在有限独立性下,大k的min-wise哈希可实现多项式误差,且种子长度与支持大小下界一致。
AI 中文摘要
最小哈希及其(k)-最小哈希变体是相似性估计、采样、草图绘制和流处理中的标准工具。一个(k)-最小哈希族要求对于固定集合的每个规定的(r)-子集(r≤k),以近似完全随机的概率出现为(r)个最小哈希值,误差在乘法因子(δ)内。先前分析表明(O(log(1/δ)+kloglog(1/δ)))-wise独立性就足够了。对于(k = Θ(log N))和(δ = N^(-c)),标准多项式构造使用(O(klog Nloglog N))个种子位。最近Chen、Huang和Li的工作对于(k = log^(O(1))N)实现了最优的(O(klog N))种子长度,但只有几乎多项式误差(2^(-O(log N/loglog N))),对于相同种子长度是否能实现多项式小误差仍未解决。我们证明标准的(s)-wise独立多项式哈希族对于[s = O(k + log(1/δ))]是(k)-最小哈希且误差为乘法因子(δ)。因此,当(k = Ω(log(1/δ)))时,只需要(O(k))-wise独立性。特别是对于(k = Θ(log N))和(δ = N^(-c)),这给出了一个明确的族,其种子长度为(O(klog N)),在常数因子上与支持大小下限匹配。证明基于规定的底部集合的条件,并仅在对由其最大哈希值给出的随机阈值进行平均后界定误差,而不是分别控制每个阈值。
英文摘要
Min-wise hashing and its $k$-min-wise variant are standard tools in similarity estimation, sampling, sketching, and streaming. A $k$-min-wise family requires every prescribed $r$-subset of a fixed set, for $r\le k$, to appear as the $r$ smallest hash values with approximately the fully random probability, up to multiplicative error $δ$. Previous analyses show that $O(\log(1/δ)+k\log\log(1/δ))$-wise independence suffices. Consequently, for $k=Θ(\log N)$ and $δ=N^{-c}$, the standard polynomial construction uses $O(k\log N\log\log N)$ seed bits. Recent work of Chen, Huang, and Li achieves the optimal $O(k\log N)$ seed length for $k=\log^{O(1)}N$, but only with almost-polynomial error $2^{-O(\log N/\log\log N)}$, leaving open whether polynomially small error is possible with the same seed length. We prove that the standard $s$-wise independent polynomial hash family is $k$-min-wise with multiplicative error $δ$ for $s=O(k+\log(1/δ)).$ Thus, when $k=Ω(\log(1/δ))$, only $O(k)$-wise independence is required. In particular, for $k=Θ(\log N)$ and $δ=N^{-c}$, this gives an explicit family with seed length $O(k\log N)$, matching the support-size lower bound up to constant factors. The proof conditions on the prescribed bottom set and bounds the error only after averaging over the random threshold given by its largest hash value, rather than controlling every threshold separately.
CommentsThis paper has been withdrawn by the authors. It has been superseded by arXiv:2607.27157, the merged version of arXiv:2607.27157v1 and arXiv:2607.10255v2