arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10500cs.DSq-bio.GN

完整标签数组的分类学分类

Taxonomic Classification with Complete Tag Arrays

Travis Gagie, Gonzalo Navarro

首次发表
浏览论文内容

中文总结 AI 辅助

提出KATKA分类器,利用游程压缩后缀数组和完整标签数组精确计数,在16S rRNA上达到93.8%属级准确率,索引更小、构建更快,优于Kraken~2和Tagger。

中文摘要 AI 辅助

诸如Kraken等分类器将参考数据库中的每个$k$-mer分配给包含它的基因组的最低共同祖先(LCA),但随着数据库的增长,这种方法效果变差,因为越来越多的$k$-mer在物种间共享。Cliffy(Ahmed, Boucher和Langmead, 2025)转而使用通过r-index找到的可变长度精确匹配,并能大致列出包含每个匹配的属;在16S rRNA上,它比Kraken~2更准确,但其索引庞大且构建成本高。我们提出KATKA,它使用Boyer--Moore--Li算法在游程编码后缀数组上查找每条read中至少给定长度的最大精确匹配(MEMs),通过完整的游程编码标签数组精确计算每个属中每个MEM的出现次数,并根据这些次数按比例给每个属计分。在SILVA 16S rRNA数据库上,KATKA的默认索引占用1.44GB,可在桌面计算机上几分钟内构建;它用单线程在66微秒内分类一条read,达到93.8%的属级准确率,接近Cliffy报告其9GB索引的准确率。对标签数组的游程进行文法压缩可将索引缩小至1.04GB,每条read耗时75微秒。在同一台机器和相同的reads上,它比Kraken~2(79.3%)和Tagger(81.7%至92.8%,取决于不一致配对的评分方式)更准确。对minimizer摘要而非序列进行索引,使索引缩小三倍,分类速度提高1.7倍,但准确率损失1.3个百分点。KATKA可在以下网址获取:此https URL。

英文摘要

Taxonomic classifiers such as Kraken assign each $k$-mer of a reference database to the lowest common ancestor (LCA) of the genomes containing it, but this works less well as databases grow, because more and more $k$-mers are shared across species. Cliffy (Ahmed, Boucher and Langmead, 2025) instead uses variable-length exact matches found with an r-index, and can list approximately the genera containing each match; on 16S rRNA it is more accurate than Kraken~2, but its index is large and expensive to build. We present KATKA, which finds the maximal exact matches (MEMs) of at least a given length in each read with Boyer--Moore--Li on a run-length compressed suffix array, counts the occurrences of each MEM in each genus exactly with a complete, run-length compressed tag array, and gives each genus credit in proportion to those counts. On the SILVA 16S rRNA database, KATKA's default index takes 1.44\,GB and can be built in minutes on a desktop computer; it classifies a read in 66\,$μ$s with one thread and reaches 93.8\% genus-level accuracy, close to what Cliffy reports for its 9\,GB index. Grammar-compressing the runs of the tag array shrinks the index to 1.04\,GB, at 75\,$μ$s per read. On the same machine and reads, it is more accurate than Kraken~2 (79.3\%) and Tagger (81.7 to 92.8\%, depending on how mates that disagree are scored). Indexing minimizer digests instead of the sequences makes the index three times smaller and classification 1.7 times faster, at a cost of 1.3 points of accuracy. KATKA is available at https://github.com/TravisGagie/KATKA.

发表机构

  • CeBiB – Center for Biotechnology and Bioengineering(CeBiB——生物技术与生物工程中心)
  • University of Chile(智利大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑