arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OpenMP5卸载的量子启发进化优化在GPU生态系统中的可移植性与性能评估

Evaluation of portability and performance of an OpenMP5 offloaded Quantum-Inspired Evolutionary Optimization Across the GPU Ecosystem

Kasturi Venkata Srikanth, Ashish Singh, Ferdin Sagai Don Bosco, Aman Mittal, Abhishek Singh, Aditya Singh, Abhishek Chopra

arXiv 2609.30862首次发表:更新:

发表机构

BosonQ Psi Corporation(BosonQ Psi 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文评估单一OpenMP5 QIEO源代码在多种GPU上的可移植性与性能,通过基因级卸载在V100、A100、MI300X上分别实现90倍、136倍、155倍于单核CPU的加速,并发现DetermineElite更适合主机执行。

AI 中文摘要

量子启发进化优化(QIEO)是一类新型的基于种群的元启发式优化算法,它将设计变量表示为一组量子比特,并通过旋转量子比特的振幅对来搜索连续的多维景观。每一代都将这些振幅朝向一个单一精英(即该代最佳个体)旋转。每代的计算成本为$O(N_p N_g)$,其中$N_p$为染色体数,$N_g$为基因(决策变量)数。此类求解器的生产使用很少局限于单一机器类别。原型通常在实验室服务器上运行,然后转移到租用的云工作站进行更复杂的活动。最大的问题则保留给领导级加速器。本文探讨是否一个单一的OpenMP 5 QIEO源代码(通过\ exttt{\#pragma omp target}卸载)能在上述每种环境中成为可行的生产路径。我们报告了针对同一源代码的多核Intel CPU基线,对0/1背包问题进行的三个独立活动。该研究包含约3000次运行,涵盖不同的染色体和基因数量,并在NVIDIA Tesla V100 SXM2、NVIDIA A100 80GB和AMD Instinct MI300X GPU上使用染色体级和基因级卸载策略进行评估。部署特有的细微差别,如Volta的常量内存悬崖、Ampere的L2持久性和\ exttt{this http URL}、CDNA 3的Infinity Cache和XCD占用率,均已处理以确保这些平台的高性能。结果显示,基因并行卸载在V100、A100和MI300X上相对于单CPU核心分别实现了90倍、136倍和155倍的几何平均加速,相对于72个主机线程分别实现了12倍、17倍和16.6倍的加速。此外,DetermineElite(即$O(N_p)$选择代最佳染色体的过程)被发现更适合在主机上执行而非设备上。

英文摘要

Quantum-inspired evolutionary optimization (QIEO) is a new class of population-based metaheuristic optimization algorithms which represents design variables as a set of qubits and searches a continuous, multi-dimensional landscape through rotation of the qubit's amplitude pair. Every generation rotates those amplitudes toward a single elite, which corresponds to that generation's best. The per-generation cost scales as $O(N_p N_g)$ for $N_p$ chromosomes and $N_g$ genes (decision variables). Production use of such solvers is rarely confined to a single machine class. Prototypes are run on laboratory servers- before moving to rented cloud workstations for more involved campaigns. The largest problems are reserved for leadership-class accelerators. This paper asks whether a \emph{single} OpenMP~5 source of QIEO, offloaded with \texttt{\#pragma omp target}, is a viable production path in each of those settings. We report three independent, campaigns of the 0/1 knapsack problem against a same-source multi-core Intel CPU baseline. The study comprises approximately 3,000 runs spanning varying chromosome and gene counts, evaluated using both chromosome-level and gene-level offload strategies on the NVIDIA Tesla V100 SXM2, NVIDIA A100 80GB, and AMD Instinct MI300X GPUs. Deployment-specific nuances such as Volta's constant-memory cliffs, Ampere's L2 persistence and \texttt{cp.async}, CDNA~3's Infinity Cache and XCD occupancy are addressed to ensure high performance of these platforms. Results reveal gene-parallel offload achieved geometric-mean speedups of 90$\times$, 136$\times$, and 155$\times$ over a single CPU core on the V100, A100, and MI300X, respectively, and 12$\times$, 17$\times$, and 16.6$\times$ over 72 host threads. Furthermore DetermineElite, the $O(N_p)$ selection of the generation-best chromosome, is found to be better suited to the host than to the device.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑