arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21233cs.AIcs.NE

面向大规模旅行商问题的广义分区交叉算法的细粒度GPU并行化

Fine-Grain GPU Parallelization of the Generalized Partition Crossover for Large-Scale Traveling Salesman Problems

Swetha Varadarajan, Darrell Whitley

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大规模旅行商问题,提出了广义分区交叉算法分区阶段的细粒度GPU并行实现,通过CUDA相关技术实现算子级并行,在1万至200万城市规模实例上取得48倍至625倍加速,提升了遗传算法求解器的可扩展性。

中文摘要 AI 辅助

旅行商问题(TSP)是研究最广泛的NP-hard优化问题之一。基于遗传算法(GA)的求解器,如边组装交叉(EAX),在许多基准实例上达到了最先进的性能。然而,这些方法在大规模并行架构中的可扩展性仍然有限,因为交叉操作涉及不规则的内存访问模式、图遍历和顺序依赖。现有的基于GPU的TSP求解器主要利用种群级并行性,且仅适用于相对较小的问题规模。本研究提出了一种针对大规模TSP实例的广义分区交叉(GPX)算子的分区阶段的细粒度GPU实现。该方法将GPX分区重新表述为图并行问题,采用合并内存布局、幽灵节点变换和连通分量分析,通过CUDA并行化父路径的合并、四度顶点的拆分、公共边的删除以及重组分量的识别。在规模从10000到200万座城市的实例上的实验结果表明,与朴素的顺序CPU实现相比,该方法实现了显著加速,GPU分区的加速比在48倍到625倍之间,同时大幅降低了内存开销。结果表明,算子级并行性可显著提高基于GA的TSP求解器在现代众核架构上的可扩展性。

英文摘要

The Traveling Salesman Problem (TSP) is one of the most extensively studied NP-hard optimization problems. Genetic Algorithm (GA)-based solvers, such as the Edge Assembly Crossover (EAX), achieve state-of-the-art performance on many benchmark instances. However, the scalability of these approaches in massively parallel architectures remains limited because crossover operations involve irregular memory access patterns, graph traversals, and sequential dependencies. Existing GPU-based TSP solvers primarily exploit population-level parallelism and are limited to relatively small problem sizes. This work presents a fine-grain GPU implementation of the partition phase of the Generalized Partition Crossover (GPX) operator for large-scale TSP instances. The proposed approach reformulates GPX partitioning as a graph-parallel problem using coalesced memory layouts, ghost-node transformations, and connected-component analysis. The im- plementation parallelizes the union of parent tours, the splitting of degree- four vertices, the deletion of common edges, and the identification of recombining components using CUDA. Experimental results on instances ranging from 10,000 to 2 million cities demonstrate substantial acceleration over a naive sequential CPU imple- mentation. The proposed GPU partitioning achieves speedups between 48x and 625x while significantly reducing memory overhead. The re- sults demonstrate that operator-level parallelism can substantially im- prove the scalability of GA-based TSP solvers on modern many-core architectures.

发表机构

  • Seattle University(西雅图大学)
  • Colorado State University(科罗拉多州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑