arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07991cs.ITmath.IT

基于组合DNA存储的错误表征与错误校正方法

Error characterization and error correction approaches in combinatorial DNA-based storage

Inbal Preuss, Omer Sabary, Ryan Gabrys, Zohar Yakhini, Eitan Yaakobi, Leon Anavy

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对组合DNA存储的不对称删除错误,开发了结合张量积码的新型纠错码,其在低测序深度场景下的解码准确率优于二维RS码。

中文摘要 AI 辅助

DNA数据存储近期已成为一种极具前景的归档解决方案,具备空间效率高、存储寿命长的优势。组合DNA编码通过DNA短序列(shortmer)的组合提升逻辑密度,其中每个序列位置由一组预定义的短DNA片段表示,可利用更少的合成周期编码更多数据,但该方法会引入独特的合成与测序错误。本研究对组合DNA存储系统中的错误进行了表征,发现不对称组合删除错误(定义为从定义组合字母的集合中遗漏单个短序列)是一种普遍存在的错误类型,尤其在读取覆盖度有限的大规模系统中更为突出。在两个已发表的数据集里,我们观察到删除错误频率很高,缺失序列会阻碍组合字母的重构。我们开展了大规模概念验证实验,确认删除错误会随测序深度降低而愈发显著:当每个序列的读取数低于50时,其频率急剧上升。针对这些错误,我们开发了一种不对称纠错码,利用张量积码将标准删除纠错码与替换纠错码(如里德-所罗门(RS)码)以及不对称Varshamov-Tenengolts码相结合。我们通过模拟和第二项大规模实验验证了其性能,将其与更简单的二维RS方案直接对比,结果显示本方法在删除主导场景中始终优于二维RS,且在二维RS难以解码数据的低覆盖度条件下,展现出更出色的解码准确率。我们的研究表明,针对组合DNA中错误的不对称特性设计纠错方案具有重要意义。

英文摘要

Data storage in DNA has recently emerged as a promising archival solution, offering space-efficient and long-lasting digital storage. Combinatorial DNA encoding enhances this potential by increasing the logical density through combinations of DNA shortmers, where each sequence position is represented by a set of predefined short DNA fragments, allowing more data to be encoded using fewer synthesis cycles. However, this method introduces unique synthesis and sequencing errors. In this study, we characterize errors in combinatorial DNA-based storage systems. We reveal that asymmetric combinatorial erasure errors, defined as the omission of a single shortmer from the set defining the combinatorial letter, are a prevalent error type, particularly in large-scale systems where read coverage is limited. In two previously published datasets, we observed a high frequency of erasure errors, where missing sequences obstruct the reconstruction of combinatorial letters. We conducted a large-scale experimental proof-of-concept and confirmed that erasure errors become increasingly prominent with reduced sequencing depth: below 50 reads per sequence, their frequency sharply increased. We developed an asymmetric error-correcting code for these errors, utilizing tensor-product codes to integrate standard erasure and substitution-correcting codes (such as Reed-Solomon (RS) codes) with asymmetric Varshamov-Tenengolts codes. We validated its performance in simulations and in a second large-scale experiment directly comparing it with the more straightforward 2D RS scheme. Our method consistently outperformed 2D RS, particularly in erasure-dominated scenarios, and demonstrated superior decoding accuracy under low coverage conditions where 2D RS struggled to decode the data. Our findings demonstrate the importance of error correction schemes tailored to the asymmetric nature of errors in combinatorial DNA.

↑