发表机构
McMaster University(麦克马斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对不确定字符串匹配,提出新的坏字符规则和快速好后缀过程,实验证明Fast_BM_Indet混合算法显著优于现有方法,并提供实用选择指南。
AI 中文摘要
我们研究不确定字符串上的精确模式匹配,其中文本或模式位置可能表示一组符号而非单个字母。聚焦于Boyer-Moore风格的方法,我们提出了新的坏字符规则(BC规则I-IV)和一种新的好后缀过程,该过程由Fast_GSR_Indet_Shift计算,通过使用单个预处理的位置索引表进行移位,避免了BM_Indet[12]中每次对齐时的重新计算。我们对十六种算法进行了系统的实验评估,包括经典的坏字符改编(如Horspool、Sunday和Zhu-Takaoka)以及将这些坏字符规则与Fast_GSR_Indet_Shift相结合的混合算法。在合成缩放实验和E. coli K-12 MG1655基因组案例研究中,Fast_BM_Indet混合算法始终优于BM_Indet和KMP_Indet[12],在某些设置下性能提升高达两个数量级。我们还发现,Zhu-Takaoka是小字母表和基因组数据上最强的仅坏字符改编算法,而使用BC规则I的Fast_BM_Indet变体提供了相当的性能,使其对较大字母表具有吸引力。最后,我们为不确定字符串应用中选择这些变体提供了实用指导。
英文摘要
We study exact pattern matching on indeterminate strings, where a text or pattern position may represent a set of symbols rather than a single letter. Focusing on Boyer-Moore-style methods, we present new bad-character rules (BC Rules I-IV) and a new good-suffix procedure, computed by Fast_GSR_Indet_Shift, which avoids the per alignment recomputation used in BM_Indet [12] by shifting with a single preprocessed position-indexed table. We conduct a systematic experimental evaluation of sixteen algorithms, including classical bad-character adaptations (e.g., Horspool, Sunday, and Zhu-Takaoka) and hybrids that combine these bad-character rules with Fast_GSR_Indet_Shift. Across synthetic scaling experiments and a case study on the E. coli K-12 MG1655 genome, the Fast_BM_Indet hybrids consistently outperform BM_Indet and KMP_Indet [12], in some settings by up to two orders of magnitude. We also find that Zhu-Takaoka is the strongest bad-character-only adaptation on small alphabets and genomic data, while the Fast_BM_Indet variant using BC Rule I offers comparable performance, making it attractive for larger alphabets. We conclude with practical guidance on choosing among these variants for indeterminate string applications.
Comments17 pages, 3 Figures, 1 Table