arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于三分搜索的快速学习型Count-Min Sketch构建

Fast Construction of Learned Count-Min Sketch via Ternary Search

Ryusuke Inami, Yusuke Matsui

arXiv 2608.08615首次发表:更新:

AI 中文总结

本研究针对学习型Count-Min Sketch参数优化低效问题,提出基于三分搜索的快速优化方法,使参数优化速度较暴力搜索提升216-740倍,且性能相当。

AI 中文摘要

学习型Count-Min Sketch(LCMS)是一种用于估算多重集中元素频率的学习型数据结构,实验表明其在容量-精度权衡上优于经典数据结构,但性能高度依赖参数选择,而此前因未充分讨论系统优化,需依赖低效的暴力搜索方法。本研究提出一种快速优化原始LCMS参数的方法,实验证实当机器学习模型性能足够好时,使用单个哈希函数即可优化加权误差指标,基于此引入三分搜索方法高效查找Unique Buckets的最优比例。该方法在模型性能充足时,性能与暴力搜索相当,参数优化速度提升216-729倍;即使模型性能欠佳,参数确定速度仍比暴力搜索快221-740倍。

英文摘要

The Learned Count-Min Sketch (LCMS) is a learned data structure that estimates element frequencies in a multiset and has been experimentally shown to outperform classical data structures in the capacity-accuracy trade-off. However, its performance depends heavily on parameter selection. Because systematic optimization has not been adequately discussed, previous approaches relied on inefficient brute-force methods. In this study, we propose a method to rapidly optimize the parameters of the original LCMS. We experimentally confirmed that when the machine learning model performs well enough, using a single hash function is sufficient to optimize the weighted error metric. Based on this, we introduce a ternary search approach to efficiently find the optimal proportion of Unique Buckets. Our method achieves the same performance as brute-force approaches while speeding up parameter optimization by $216$-$729$ times when the machine learning model's performance is sufficient. Furthermore, even when the model's performance is suboptimal, our approach still determines appropriate parameters $221$-$740$ times faster than the brute-force approach.

CommentsSISAP short paper (full version)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑