发表机构
Singapore University of Technology and Design; Institute of Molecular and Cell Biology (IMCB), Agency for Science, Technology and Research (A*STAR); RIKEN Center for Computational Science(新加坡科技设计大学; 新加坡科技研究局分子与细胞生物学研究所(IMCB); 理化学研究所计算科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对癌症多基因组合识别计算难题,提出基于对偶剪枝的深度优先搜索定价算法,在单CPU上秒级求解,实现首个保证全局最优的精确方法。
AI 中文摘要
癌症由估计2至9个基因突变驱动,这些突变被称为多基因组合。由于基因数量庞大,识别这些多基因组合构成一个计算上具有挑战性的问题,此前需要超级计算机枚举所有或大部分基因组合。近期工作将该分类任务表述为多基因癌症驱动集合覆盖问题(MHCDSCP),并通过在单台普通CPU上使用列生成方法求解,其中定价问题识别有前景的基因组合(arXiv:2602.22551)。主要瓶颈在于需要反复求解的定价问题。为最优求解定价问题,我们提出一种专用的深度优先搜索方法,枚举所有可能的基因组合。尽管使用深度优先搜索枚举基因组合此前已在约150,000个CPU核心的超级计算机上执行(arXiv:2603.16721),我们的深度优先搜索算法通过利用对偶信息剪枝搜索空间的大部分,在单台CPU上数秒内即可最优求解定价问题。通过消除这一瓶颈,我们的方法首次为MHCDSCP提供精确方法,保证找到全局最优解。
英文摘要
Cancer is driven by an estimated two to nine gene mutations, known as multi-hit combinations. Due to the large number of genes, identifying these multi-hit combinations presents a computationally challenging problem, which has previously required supercomputers to enumerate all or most gene combinations. Recent work formulated this classification task as the Multi-Hit Cancer Driver Set Cover Problem (MHCDSCP) and solved it via column generation on a single commodity CPU, where a pricing problem identifies promising gene combinations (arXiv:2602.22551). The main bottleneck is the pricing problem that has to be solved repeatedly. To solve the pricing problem to optimality, we propose a dedicated depth first search method that enumerates all possible gene combinations. Although enumerating gene combinations with depth first search have been performed on supercomputers with around 150,000 CPU cores (arXiv:2603.16721), our depth first search algorithm solves the pricing problem to optimality within seconds on a single CPU by leveraging dual information to prune large parts of the search space. By eliminating this bottleneck, our approach yields the first exact method for the MHCDSCP, which is guaranteed to find a global optimal solution.
Comments13 pages