arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MerKurio:基于k-mer匹配的序列提取与注释

MerKurio: sequence extraction and annotation based on matching k-mers

Lukas Schönmann, Heinz Himmelbauer, Juliane C. Dohm

arXiv 2609.39830首次发表:更新:

发表机构

Institute of Computational Biology, Department of Biotechnology and Food Science, BOKU University(维也纳应用自然科学大学生物技术与食品科学学院计算生物学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MerKurio是一个用Rust编写的高性能命令行工具,基于k-mer匹配实现序列提取与注释,支持双端读取和反向互补搜索,速度优于现有工具。

AI 中文摘要

k-mer,即长度为k的短子序列,在使用生物序列数据进行计算分析中扮演着重要角色。一个关键的下游处理步骤涉及回到所选k-mer的源序列以进行进一步分析或验证。在此,我们介绍MerKurio,一个用Rust编写的高性能命令行工具,具有两个主要功能:1)使用k-mer从FASTA/FASTQ文件中提取序列记录;2)使用k-mer标签对SAM/BAM文件中的比对序列进行注释或过滤。该工具通过基于查询特征的经验推导规则自动选择模式匹配算法,并支持双端读取、压缩输入文件和反向互补搜索。基准比较表明,MerKurio在速度上优于现有工具,同时提供额外功能,包括详细的匹配统计和全面的文件格式支持。MerKurio是一个基于k-mer的快速且用户友好的序列记录提取和比对序列标记工具。MerKurio可从此https URL免费获取。

英文摘要

K-mers, short subsequences of length k, play an important role in computational analyses using biological sequence data. A critical downstream processing step involves getting back to the source sequences of selected k-mers for further analysis or validation. Here, we present MerKurio, a high-performance command line tool written in Rust with two main functionalities: 1) extracting sequence records from FASTA/FASTQ files using k-mers and 2) annotating or filtering aligned sequences in SAM/BAM files with k-mer tags. The tool automatically selects the pattern matching algorithm via empirically derived rules based on query characteristics and supports paired-end reads, compressed input files, and reverse complement searching. Benchmark comparisons demonstrate that MerKurio outperforms existing tools in terms of speed while providing additional utility including detailed matching statistics and comprehensive file format support. MerKurio is a fast and user-friendly tool for sequence record extraction and tagging of aligned sequences based on k-mers. MerKurio is freely available at https://github.com/lschoenm/MerKurio.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑