发表机构
Technion—Israel Institute of Technology; Eindhoven University of Technology(以色列理工学院; 埃因霍温理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对DNA存储中的覆盖深度问题,推导期望公式,证明q元单纯码和汉明码的最优性,并分析随机线性码与MDS基准的差距有界。
AI 中文摘要
DNA存储系统通过随机采样合成的DNA链来检索信息,使得成功恢复所需的读取次数成为基本的性能度量。这引出了覆盖深度问题:对于给定的码参数,确定一个线性码,使其最小化随机采样列以张成整个信息空间的期望次数。虽然MDS码在存在时已知是最优的,但在MDS码不存在的参数范围内,确定最优码在很大程度上仍是未解决的问题。在这项工作中,我们推导了覆盖深度设置中期望值的一般公式,并应用该结果证明了q元单纯码在其存在参数下的最优性。此外,通过将覆盖深度与对偶码中的独立集联系起来,并利用拟阵基生成多项式的对数凹性,我们证明了q元汉明码的最优性。我们进一步分析了随机线性码,推导出其期望覆盖深度的精确表达式,并表明在恒定速率范围内,它们与MDS基准的加性差距保持有界,且与码长无关。
英文摘要
DNA storage systems retrieve information by randomly sampling synthesized DNA strands, making the number of reads required for successful recovery a fundamental performance measure. This motivates the coverage depth problem: for given code parameters, determine a linear code that minimizes the expected number of randomly sampled columns required to span the entire information space. While MDS codes are known to be optimal whenever they exist, identifying optimal codes in parameter regimes where MDS codes do not exist remains largely open. In this work, we derive a general formula for the expectation in the coverage depth setting, and apply this result to establish the optimality of the $q$-ary simplex code for the parameters that allow its existence. Moreover, we show the optimality of the $q$-ary Hamming code by relating coverage depth to independent sets in the dual code and exploiting log-concavity properties of matroid basis-generating polynomials. We further analyze random linear codes, derive an exact expression for their expected coverage depth, and show that, in the constant-rate regime, their additive gap from the MDS benchmark remains bounded independently of the code length.