arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分布式DNA数据存储中的覆盖深度问题

The Coverage Depth Problem in Distributed DNA Data Storage

Xiangliang Kong, Ohad Elishco, Chen Wang, Tolga M. Duman

arXiv 2610.02931首次发表:更新:

发表机构

State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, Chinese Academy of Sciences; School of Electrical and Computer Engineering, Ben-Gurion University of the Negev; Department of Computer Science, Technion – Israel Institute of Technology; Department of Electrical and Electronics Engineering, Bilkent University(中国科学院数学与系统科学研究院; 内盖夫本古里安大学电气与计算机工程学院; 以色列理工学院计算机科学系; 比尔肯特大学电气与电子工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究分布式DNA数据存储中的覆盖深度问题,推导恢复时间与期望读取成本的精确公式,证明MDS码的最优性及单纯码的唯一最小化性质,并阐明分布式采样对延迟和测序成本的影响。

AI 中文摘要

DNA测序中的随机采样会产生重复读取,增加检索延迟和测序成本。我们研究了在无噪声均匀采样下分布式DNA存储中全消息恢复的覆盖深度问题,其中链被分配到$M$个容器中,每轮从每个容器中独立地有放回地采样一条链。对于任意线性码和有序划分,我们推导了恢复时间分布和期望的精确公式。我们证明了MDS码(若存在)对每个固定划分都是最优的,并建立了期望总读取成本的通用下界及其等式条件。对于MDS码,我们确定了能带来总读取量真正节省的容器大小区间,以及仅提供并行性而不改变渐近测序成本的区间。对于单纯码,我们证明了$q$元单纯码在同构意义上是具有相同参数的码中唯一的单容器最小化器,解决了Bertuzzo、Ravagnani和Yaakobi最近的猜想。我们进一步构造了一个达到最小总读取成本的划分,并推导了中间和平衡划分的界。这些结果阐明了分布式采样何时仅减少延迟,以及何时也降低测序成本。

英文摘要

Random sampling in DNA sequencing produces repeated reads, increasing retrieval latency and sequencing cost. We study the coverage-depth problem for full-message recovery in distributed DNA storage under noiseless uniform sampling, where strands are partitioned among $M$ containers and one strand is independently sampled with replacement from each container per round. For arbitrary linear codes and ordered partitions, we derive exact formulas for the recovery-time distribution and expectation. We prove that MDS codes, whenever they exist, are optimal for every fixed partition, and establish a universal lower bound on the expected total read cost together with its equality conditions. For MDS codes, we identify container-size regimes that yield genuine savings in total reads and regimes that provide only parallelism without changing the asymptotic sequencing cost. For simplex codes, we prove that the $q$-ary simplex code is, up to isomorphism, the unique single-container minimizer among codes with the same parameters, resolving a recent conjecture by Bertuzzo, Ravagnani, and Yaakobi. We further construct a partition attaining the minimum total read cost and derive bounds for intermediate and balanced partitions. These results clarify when distributed sampling reduces latency alone and when it also reduces sequencing cost.

Comments33 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑