arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31501cs.ITeess.SPmath.IT

免疫基因组学中序列重建问题的基本极限

Fundamental Limits of Sequence Reconstruction Problems in Immunogenomics

Jaswanthi Mandalapu, Deepak Charan, V. Arvind Rameshwar, Nir Weinberger

AI总结:

本研究针对免疫基因组学中D基因重建问题,确定了三种痕迹生成模型下的最优痕迹复杂度,并开发了相应的低复杂度解码算法,填补了原始工作的空白。

AI中文摘要:

个性化免疫基因组学的目标是从经过修剪、延伸和突变改变的“库序列”中恢复个体的种系免疫球蛋白基因片段。在本工作中,我们研究了由Bhardwaj等人(2021)引入的三种生物学动机下的痕迹生成模型下D基因重建的基本痕迹复杂度(即准确重建所需的样本/痕迹数量),并为原始工作中遗留的重建问题开发了实用算法。首先,对于TrimSuffixAndExtend模型,我们确定了最优痕迹复杂度为{\Theta(n)},并开发了一种低复杂度的前缀过滤模式(PFM)解码器,该解码器实现了这一复杂度规模。其次,对于密切相关的双侧TrimAndExtend模型,我们展示了最优痕迹复杂度反而是{\Theta(n^2)},并由简单的按位模式(BWM)解码器实现。第三,对于SuffixExtend-t(TrimSuffix)模型,我们建立了痕迹复杂度的多项式分离的下界和上界。我们的结果源于信息论下界与所提出的重建算法的严密分析的结合。

英文摘要:

The goal of personalized immunogenomics is to recover an individual's germline immunoglobulin gene segments from 'repertoire sequences' altered by trimming, extension, and mutation. In this work, we study the fundamental trace complexity (or the number of samples/traces required for accurate reconstruction) of D-gene reconstruction under three biologically motivated trace-generation models introduced by Bhardwaj et al. (2021) and develop practical algorithms for reconstruction problems left open in the original work. First, for the TrimSuffixAndExtend model, we establish the optimal trace complexity to be {Θ(n)}, and develop a low-complexity Prefix-Filtered Mode (PFM) decoder that achieves this scaling. Second, for the closely related two-sided TrimAndExtend model, we show that the optimal trace complexity is instead {Θ(n^2)}, and is achieved by the simple Bit-Wise Mode (BWM) decoder. Third, for the SuffixExtend-t(TrimSuffix) model, we establish polynomially separated lower and upper bounds on trace complexity. Our results follow from information-theoretic lower bounds coupled with tight analyses of the proposed reconstruction algorithms.

↑