arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17353cs.CC

面向最优前缀无关图构建:NP-困难性与结构洞察

Towards Optimal Prefix-Free Graph Construction: NP-Hardness and Structural Insights

  • Institute of Clinical and Translational Research, Biomedical Research Center of the Slovak Academy of Sciences(斯洛伐克科学院生物医学研究中心临床与转化研究所)
  • National Institute for Research and Development in Informatics – ICI Bucharest(布加勒斯特国家信息与研发研究所)

机构由 AI 辅助整理,请以论文原文为准。

Andrej Baláž, Alexandru Popa

AI总结:

本研究证明最小前缀无关图构建是NP-困难的,并建立其与de Bruijn图的结构联系,提出O(2^q n)固定参数算法,为重复泛基因组紧凑表示奠定理论基础。

AI中文摘要:

前缀无关解析为构建大型且重复的泛基因组压缩表示提供了一种高效方法,并自然诱导出一种称为前缀无关图的图表示。在本工作中,我们首次对构建最小尺寸前缀无关图的问题展开理论研究,其中尺寸同时考虑不同片段标签的总长度以及表示输入序列的路径。我们证明,即使触发器由单个字符组成,选择最优触发器集合也是NP-困难的。利用同步码归约,我们将此困难性结果推广到任意固定触发器长度,并进一步证明该问题在三字母表上仍为NP-困难。随后,我们建立了前缀无关图与de Bruijn图之间的结构联系。特别地,我们证明每个压缩de Bruijn图都可以实现为前缀无关图,并推导出最小泛基因组图、最小前缀无关图、压缩de Bruijn图和de Bruijn图之间尺寸关系的层级。最后,我们给出一个精确的固定参数算法,运行时间为$O(2^q n)$,其中$q$是不同候选触发器词的数量,$n$是泛基因组总长度。我们的结果刻画了优化前缀无关图表示的计算局限性和结构性质,并为设计重复泛基因组数据的紧凑图表示提供了理论基础。

英文摘要:

Prefix-free parsing provides an efficient way to construct compressed representations of large and repetitive pangenomes and naturally induces a graph representation known as a prefix-free graph. In this work, we initiate a theoretical study of the problem of constructing prefix-free graphs of minimum size, where the size accounts for both the total length of distinct segment labels and the paths representing the input sequences. We show that selecting an optimal set of trigger words is NP-hard, already when triggers consist of single characters. Using a synchronized-code reduction, we extend this hardness result to every fixed trigger length and further show that the problem remains NP-hard over an alphabet of size three. We then establish a structural connection between prefix-free graphs and de Bruijn graphs. In particular, we show that every compacted de Bruijn graph can be realized as a prefix-free graph and derive a hierarchy relating the sizes of minimum pangenomic graphs, minimum prefix-free graphs, compacted de Bruijn graphs, and de Bruijn graphs. Finally, we give an exact fixed-parameter algorithm running in $O(2^q n)$ time, where $q$ is the number of distinct candidate trigger words and $n$ is the total pangenome length. Our results characterize both the computational limitations and the structural properties of optimizing prefix-free graph representations and provide a theoretical foundation for the design of compact graph representations of repetitive pangenomic data.

补充信息

↑