arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以存储为中心的系统设计,用于实现快速、高效且低成本的基因组与宏基因组分析

Storage-Centric System Designs for Enabling Fast, Efficient, and Low-Cost Genomic and Metagenomic Analyses

Nika Mansouri Ghiasi

arXiv 2608.31004首次发表:更新:

AI 中文总结

本论文针对基因组与宏基因组分析的数据移动及准备瓶颈,提出定制化以存储为中心的系统,可在存储内分析数据、实现高压缩存储与高性能访问,显著提升分析的性能、能效及成本效益。

AI 中文摘要

基因组与宏基因组分析在精准医学、紧急临床场景、传染病早期预警发现、通过病原体监测保障食品安全、农业及科学发现等众多领域发挥关键作用。由于分析和存储海量基因组与宏基因组序列数据面临挑战,学界已投入大量精力加速(宏)基因组分析并对序列数据进行压缩存储。尽管这些技术具有优势,但我们在访问存储的序列数据并将其提供给分析单元时,发现两个主要未解决的问题:(i)数据移动瓶颈,即需从存储中移动大量低复用数据,给系统其他部分造成不必要负担;(ii)数据准备瓶颈,即压缩后的序列数据在分析前需先解压并格式化。在本论文中,我们提出定制化的以存储为中心的系统,该系统可高效实现两项功能:(i)在存储系统内部分析(宏)基因组数据;(ii)实现大规模序列数据的高压缩存储与高性能访问,从而缓解数据移动、计算及数据准备的开销。我们验证了所提出的系统可显著提升(宏)基因组分析的系统性能、能效及成本效益。我们希望本论文提出的以存储为中心的系统能推动(宏)基因组分析的更广泛应用,并启发未来研究从根本上改善健康与生命科学相关其他数据密集型应用领域的性能、能效及成本效益。

英文摘要

Genomic and metagenomic analyses play critical roles in many fields, such as precision medicine, urgent clinical settings, discovering early warnings of communicable diseases, ensuring food safety through pathogen monitoring, agriculture, and scientific discovery. Due to the challenges of analyzing and storing massive volumes of genomic and metagenomic sequence data, significant efforts have been made to accelerate (meta)genomic analyses and store sequence data compressed. Despite the benefits of these techniques, we identify two major outstanding problems in accessing stored sequence data and supplying it to the analysis units: (i) the data movement bottleneck due to moving large amounts of low-reuse data from storage and the unnecessary burden on the rest of the system, and (ii) the data preparation bottleneck, where compressed sequence data needs to be first decompressed and formatted before analysis. In this dissertation, we present customized storage-centric systems, which efficiently (i) analyze (meta)genomic data inside the storage system, and (ii) enable highly-compressed storage and high-performance access of large-scale sequence data, thereby alleviating the overheads of data movement, computation, and data preparation. We demonstrate that the proposed systems significantly improve system performance, energy efficiency, and system cost-efficiency of (meta)genomic analysis. We hope that the storage-centric systems proposed in this dissertation facilitate the broader adoption of (meta)genomic analyses and inspire future research to fundamentally improve the performance, energy efficiency, and cost-effectiveness of other data-intensive application domains related to health and life sciences.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑