AI 中文总结
本研究针对学习型Count-Min Sketch参数优化低效问题,提出基于三分搜索的快速优化方法,使参数优化速度较暴力搜索提升216-740倍,且性能相当。
AI 中文摘要
学习型Count-Min Sketch(LCMS)是一种用于估算多重集中元素频率的学习型数据结构,实验表明其在容量-精度权衡上优于经典数据结构,但性能高度依赖参数选择,而此前因未充分讨论系统优化,需依赖低效的暴力搜索方法。本研究提出一种快速优化原始LCMS参数的方法,实验证实当机器学习模型性能足够好时,使用单个哈希函数即可优化加权误差指标,基于此引入三分搜索方法高效查找Unique Buckets的最优比例。该方法在模型性能充足时,性能与暴力搜索相当,参数优化速度提升216-729倍;即使模型性能欠佳,参数确定速度仍比暴力搜索快221-740倍。
英文摘要
The Learned Count-Min Sketch (LCMS) is a learned data structure that estimates element frequencies in a multiset and has been experimentally shown to outperform classical data structures in the capacity-accuracy trade-off. However, its performance depends heavily on parameter selection. Because systematic optimization has not been adequately discussed, previous approaches relied on inefficient brute-force methods. In this study, we propose a method to rapidly optimize the parameters of the original LCMS. We experimentally confirmed that when the machine learning model performs well enough, using a single hash function is sufficient to optimize the weighted error metric. Based on this, we introduce a ternary search approach to efficiently find the optimal proportion of Unique Buckets. Our method achieves the same performance as brute-force approaches while speeding up parameter optimization by $216$-$729$ times when the machine learning model's performance is sufficient. Furthermore, even when the model's performance is suboptimal, our approach still determines appropriate parameters $221$-$740$ times faster than the brute-force approach.
CommentsSISAP short paper (full version)