MultiTable: 在任意物理负载因子直至并包括1下更快的哈希表
MultiTable: A Faster Hash Table at any Physical Load Factor up to and Including One
AI总结:
提出多表哈希表,在任意负载因子下性能优于 hashbrown,吞吐量提升至2.1倍,内存节省最多22%,且无需重新哈希即可增长。
AI中文摘要:
我们提出 multitable 及其 Rust 参考实现:一种稳定的哈希表,在相同物理内存下速度显著更快,并且比 Rust 的 hashbrown 实现中的 SwissTable 更灵活。作为 84 种配置的算术平均值,当两者对相同原始字节进行哈希时,它提供 hashbrown 吞吐量的 2.1 倍;当 hashbrown 以原生整数为键(其最佳情况)时,提供 1.9 倍;仅负查找时,分别为 3.2 倍和 2.9 倍。Multitable 可达到任意物理负载因子直至并包括 1(演示了 0.9999),恰好满足请求的容量,而 hashbrown 在 4 字节键和值下于 0.777 处加倍。在 hashbrown 75% 饱和度(其刚性阶梯的假设平均情况)和 multitable 尺寸为 0.97 物理负载因子时,hashbrown 占用空间多 66%。查找探测次数在负载因子接近 1 时没有悬崖。桶大小、物理负载因子和失败预算是参数,multitable 可以在不重新哈希的情况下增长。我们实现了 multitable 的两种变体:普通和过滤。在 Apple M2 Pro 上相同物理内存下,过滤 multitable 在所有 84 种插入、命中和未命中配置中领先 hashbrown。Multitable 更节省内存,在 4 字节键和值的映射上相同混合查找吞吐量下,过滤 multitable 比 hashbrown 少需要最多 12% 的字节,普通 multitable 小 18%,在相同内存中容纳多 22% 的键。
英文摘要:
We present \emph{multitable} and its Rust reference implementation: a stable hash table both materially faster at equal physical memory and more flexible than the SwissTable in its Rust's hashbrown implementation. As an arithmetic mean over 84 configurations it delivers $\mathbf{2.1\times}$ hashbrown's throughput when both hash the same raw bytes and $\mathbf{1.9\times}$ when hashbrown is keyed on native integers, its best case; on negative lookups alone, $3.2\times$ and $2.9\times$. Multitable reaches \textbf{any physical load factor} up to and \textbf{including one} ($0.9999$ demonstrated), exactly for the requested capacity, compared to hashbrown which doubles at $0.777$ for 4-byte keys and values. At $75\%$ saturation of hashbrown (assumed average case of its rigid ladder) and multitable sized to $0.97$ physical load factor, hashbrown takes $66\%$ more space. The lookup probe count has no cliff as the load factor approaches one. Bucket size, physical load factor, and failure budget are parameters, and the multitable can be grown without rehashing. We implement two variants of multitable: plain and filtered. At equal physical memory on an Apple M2 Pro the filtered multitable leads hashbrown in all $84$ insert, hit, and miss configurations. Multitable is more \textbf{memory-efficient}, at equal mixed-lookup throughput on the map of $4$-byte keys and values the filtered multitable needs up to $12\%$ fewer bytes than hashbrown, and the plain multitable is $18\%$ smaller, holding $\mathbf{22\%}$ more keys in the same memory.