发表机构
Guangzhou Institute of Technology, Xidian University; State Key Laboratory of Integrated Service Networks, Xidian University(西安电子科技大学广州研究院; 西安电子科技大学综合业务网理论及关键技术国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过将DNA标记建模为确定性信道,证明其零错误容量等于标记容量,推导出所有星图的零错误容量,从而完整刻画单标签设置下的标记容量,并给出容量可达码的构造方法及容量范围。
AI 中文摘要
DNA标记在生物医学应用中日益受到关注,包括分子成像、诊断和基因组分析。在DNA标记过程中,根据特定应用的需求设计一组DNA序列模式,称为标签。对于每条DNA序列,标记过程生成一个输出序列,记录标签的位置。因此,具有不同标记输出的DNA序列可以通过标记过程加以区分。为了量化这种能力,标记容量被定义为随着序列长度趋于无穷大,通过标记过程可区分的DNA序列的最大数量的指数增长率[2]。迄今为止,在单标签设置下,若干情况的标记容量已被确定。在本文中,我们将标记过程建模为一个确定性信道,并证明其零错误容量等于标记容量。对于单个标签,相应的信道可以用星图表示。因此,刻画单个标签的标记容量等价于确定相应星图的零错误容量。我们推导了所有星图的零错误容量,从而为所有单标签情况提供了标记容量的完整刻画。此外,我们开发了一种构造容量可达码的通用方法。这些结果适用于任意有限字母表上的标记问题,并不仅限于DNA字母表。最后,对于固定的标签长度,我们精确刻画了可达标记容量的范围,并识别出达到最小和最大容量的所有单标签结构。
英文摘要
DNA labeling has attracted increasing attention in biomedical applications, including molecular imaging, diagnostics, and genomic analysis. In a DNA labeling process, a set of DNA sequence patterns, referred to as labels, is designed according to the requirements of a specific application. For each DNA sequence, the labeling process generates an output sequence that records the positions of the labels. DNA sequences with different labeling outputs can therefore be distinguished through the labeling process. To quantify this capability, the labeling capacity is defined as the exponential growth rate of the maximum number of DNA sequences that can be distinguished through the labeling process as the sequence length tends to infinity [2]. To date, the labeling capacities of several cases in the single-label setting have been determined. In this paper, we formulate the labeling process as a deterministic channel and show that its zero-error capacity is equal to the labeling capacity. For a single label, the corresponding channel can be represented by a star graph. Thus, characterizing the labeling capacity of a single label is equivalent to determining the zero-error capacity of the corresponding star graph. We derive the zero-error capacities of all star graphs, thereby providing a complete characterization of the labeling capacities for all single-label cases. Furthermore, we develop a general method for constructing capacity-achieving codes. These results apply to labeling problems over arbitrary finite alphabets and are not restricted to the DNA alphabet. Finally, for a fixed label length, we exactly characterize the range of achievable labeling capacities and identify all single-label structures that attain the minimum and maximum capacities.