arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时将哈希表桶树化:对列表、混合结构和红黑树链化的可复现 C 语言研究

When to Treeify Hash Table Buckets: A Reproducible C Study of List, Hybrid, and Red-Black Tree Chaining

Georgii Kashintsev

arXiv 2607.26530首次发表:更新:

AI 中文总结

该研究针对 C 语言哈希表,对比列表、混合结构、红黑树链化策略,发现混合增量转换性能更优,转换开销远超阈值选择,树化内存更高,压力测试下性能提升显著。

AI 中文摘要

从业者摘要:请勿仅照搬 Java 的 8 个阈值:当哈希桶长度增长时,混合批量转换(加载后转换)在插入期间仍会遍历列表,而混合增量转换(桶达到 k 时立即转换)的性能与始终树化相当。对于过载的桶,应选择混合增量转换或始终树化;当调整大小后链仍较短时,混合批量转换可保留用于纯批量加载后查询。混合增量转换近似 Java 的转换时机,而非 HashMap 的移植版本。以下核心指标为字符串比较(strcmp)次数和堆内存,比长列表的挂钟时间更稳定。当单个哈希桶变长时,链表分离链化会产生线性的每桶成本。我们表明,对于 C 语言实现者而言,转换运行(混合批量转换的最终化与混合增量转换)的开销远大于阈值 k 的选择。使用一个 C 分离链化 API,我们在均匀哈希 FNV 下(包括 alpha≈122 时的固定 m 探测)、强制桶链化压力测试,以及中等负载的同 API 规模运行(alpha=16)下比较了不同策略。压力测试下,列表查找平均约 31250 次比较,树化后约 15 次;中等负载探测下,混合批量转换需约 3700 万次比较,混合增量转换仅需约 4.6 万次;加载后最终比较结果相近(约 15 次)。树化桶的堆内存使用量约为列表的 1.7 倍。长列表的压力测试挂钟时间具有说明性且运行噪声大,因此我们重点呈现比较结果和内存数据。通过相同策略复现真实三元词组发布列表长度,得到的排名一致。在 alpha≈122 且无调整大小时,部分树化的优势实际源于延迟重哈希——当 m 过小时应先调整大小。

英文摘要

Practitioner summary. Do not copy Java's threshold of eight alone: when bins grow long, hybrid-batch (convert after load) still walks lists during insert, while hybrid-incremental (convert as soon as a bin hits k) matches always-tree. Prefer hybrid-incremental or always-tree for overloaded bins; reserve hybrid-batch for pure bulk load then query when chains stay short after resize. Hybrid-incremental approximates Java conversion timing, not a HashMap port. Lead metrics below are strcmp counts and heap - more stable than long-list wall-clock. When individual hash buckets grow long, linked-list separate chaining incurs linear per-bucket cost. We show that when conversion runs (hybrid-batch finalize vs. hybrid-incremental) dwarfs the choice of threshold k for C implementers. Using one C separate-chaining API, we compare policies under uniform-hash FNV (including a fixed-m probe at alpha ~ 122), forced-bucket chaining stress, and a moderate-load same-API scale run (alpha = 16). Under stress, list lookup averages ~31,250 comparisons vs ~15 once treeified; mid-load probes need ~37M comparisons under hybrid-batch vs ~46k under hybrid-incremental; final post-load comparisons converge (~15). Tree buckets use about 1.7x more heap than lists. Stress wall-clock for long lists is illustrative and run-noisy; we therefore headline comparisons and memory. Replaying real trigram posting-list lengths through the same policies yields the same ranking. At alpha ~ 122 without resize, some tree wins are really deferred rehash - resize first when m is simply too small.

Comments14 pages, tables in text. Intended submission: Software: Practice and Experience (Wiley). Source: https://github.com/rbtreechainingforhashtable/project (tag v0.0.1)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑