发表机构
Faculty of Mathematics and Computer Science, University of Bucharest(布加勒斯特大学数学与计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究易位事件驱动的基因组重排,证明非均匀连续与非连续易位距离问题在任意有限字母表上均为NP困难,并对单目标字符串情形给出以目标长度为参数的固定参数可处理算法。
AI 中文摘要
在本文中,我们研究由易位事件进行的基因组重排。自1936年(Dobzhansky和Sturtevant)以来,基因组重排被用于衡量生物之间的进化距离。染色体被表示为DNA字符串,\textbf{易位操作}定义为两个字符串之间前缀的交换。该操作产生两个新字符串(染色体),它们随后可用于后续的易位。如果新字符串以单拷贝形式产生,则称易位为\textbf{连续的},因此每个字符串只能用于一次后续操作。当易位操作产生的词被认为具有无限数量的拷贝时,该易位被称为\textbf{非连续的}。如果交换的前缀长度相等,则易位称为\textbf{均匀的}。否则,易位称为\textbf{非均匀的}。两个字符串集合(称为输入集合和目标集合)之间的\textbf{易位距离}表示通过易位操作获得目标集合中所有字符串所需的最少易位次数。我们证明了非均匀连续和非均匀非连续易位距离问题在任意有限字母表上都是NP困难的,其中字母表是输入的一部分。对于目标集合由单个字符串组成的情况,我们给出了一个以目标字符串长度为参数的固定参数可处理算法。
英文摘要
In this paper we study the genome rearrangements done by translocation events. Genome rearrangements were used to measure evolutionary distance between organisms since 1936 (Dobzhansky and Sturtevant). The chromosomes are represented as strings of DNA and the \emph{translocation operation} is defined as the exchange of prefixes between two strings. This operation results in the creation of two new strings (chromosomes) that can then be utilized in subsequent translocations. A translocation is referred to as \emph{contiguous} if the new strings are produced in a single copy, so each of them can be used in only one subsequent operation. When the words produced by a translocation operation are considered to have an infinite number of copies, the translocation is referred to as \emph{non-contiguous}. If the exchanged prefixes are of equal length, the translocation is called \emph{uniform}. Otherwise, the translocation is termed \emph{non-uniform}. The \emph{translocation distance} between two sets of strings, termed the input set and the target set, represents the minimum number of translocations necessary to obtain all the strings in the target set via translocation operations. We prove that both the non-uniform contiguous and the non-uniform non-contiguous translocation distance problems are NP-hard over arbitrary finite alphabets, where the alphabet is part of the input. For the case in which the target set consists of a single string, we give a fixed-parameter tractable algorithm parameterized by the length of the target string.