发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对群体遗传学中基因型图编辑的突变重新映射问题,提出批量共享逆拓扑遍历算法,实现比独立方法快10.5倍的编辑效果,且保留精确语义。
AI 中文摘要
通过插入或替换节点更新图,同时保留语义并复用现有结构,是一个反复出现的计算问题。在群体遗传学中,该问题出现在基因型表示图(GRG)中,这是一种有向无环图,通过为单个突变共享子图结构,无损编码数十万个样本的分阶段遗传变异。在GRG中,每个突变的携带者集合被隐式编码为从其分配节点可达的叶节点集合,因此更新突变是一个结构编辑问题,而当前方法会单独重新映射突变。本文提出一种批量突变重新映射算法,将独立的感知复用遍历替换为单次共享的逆拓扑遍历,一次性识别整个批次的复用候选;该遍历传播紧凑的位并行突变状态,并使用自适应的稀疏/密集携带者集合表示,覆盖从稀有到常见的变异密度。批量处理是基于拆分的并行性的内存可扩展补充,后者会为每个工作者复制图和遍历状态。我们在受控更新工作负载和端到端等位基因极化(群体遗传分析中常见的批量携带者集合更新)上评估了重新映射方法。我们的方法比独立重新映射快达10.5倍,同时保留精确的携带者集合语义。
英文摘要
Updating a graph by inserting or replacing nodes while preserving semantics and reusing existing structure is a recurring computational problem. In population genetics, this problem arises in the genotype representation graph (GRG), a directed acyclic graph that losslessly encodes phased genetic variation across hundreds of thousands of samples by sharing subgraph structure for individual mutations. In a GRG, each mutation's carrier set is implicitly encoded as the set of leaf nodes reachable from the node it is assigned to. Updating a mutation is therefore a structural editing problem, and current approaches remap mutations individually. This paper introduces a batched mutation-remapping algorithm that replaces independent reuse-aware traversals with a single shared reverse-topological pass, identifying reuse candidates for an entire batch at once. The pass propagates compact bit-parallel per-mutation state and uses an adaptive sparse/dense carrier set representation spanning rare-to-common variant densities. Batching is the memory-scalable complement to split-based parallelism, which instead replicates graph and traversal state per worker. Our remapping is evaluated on a controlled update workload and on end-to-end allele polarization, a bulk carrier set update that is common in population genetic analysis. Our approach is up to 10.5$\times$ faster than independent remapping while preserving exact carrier-set semantics.