arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07862cs.DS

并行化阶乘空间:通过双通道AVX2执行实现Steinhaus-Johnson-Trotter算法的3倍SIMD加速

Parallelizing the Factorial Space: Multi-Core OpenMP Scaling and Scalable SIMD Acceleration of the Steinhaus-Johnson-Trotter Algorithm via Dual-Lane AVX2 Execution

Serge Melnikov

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于AVX2双通道SIMD的Steinhaus-Johnson-Trotter排列生成加速框架,通过组合空间划分和向量混洗实现3倍吞吐提升,并避免存储转发停顿。

中文摘要 AI 辅助

本文提出了一种针对Steinhaus-Johnson-Trotter排列生成算法的高性能SIMD加速框架,目标平台为使用AVX2指令集的现代x86-64架构。通过利用一种新颖的组合空间划分方法,结合预计算的索引偏移量和单周期向量字节混洗操作(_mm256_shuffle_epi8),我们的双通道向量化实现在统一的执行掩码下,于单个256位YMM寄存器中处理两个独立并发的排列流。实验评估表明,与Donald Knuth的算法P(TAOCP第4A卷)以及Yusheng Hu最近的Ring-Cascade算法相比,我们的性能吞吐量提升了3倍,其中算法P此前已通过标量代码实现了3倍加速。所提出的软件架构保持了跨编译器兼容性,在热循环中完全避免了存储转发内存停顿,并已针对n=11阶进行了验证,在原生硬件上n=13的基准性能约为12.7亿CPU周期。

英文摘要

This paper presents a high-performance SIMD acceleration framework for the Steinhaus-Johnson-Trotter algorithm, targeted at modern x86-64 architectures using the AVX2 instruction set. By exploiting a novel combinatorial space partitioning combined with single-cycle vector byte shuffling ($\texttt{\_mm256\_shuffle\_epi8}$), our dual-lane vectorized implementation processes two independent, concurrent permutation streams within a single 256-bit YMM register under a uniform execution mask. Empirical evaluations demonstrate a $3\times$ throughput increase over an optimized scalar baseline of Knuth's Algorithm P (accelerated by $3\times$ via isolated sweeping branches) and outperform the recent Ring-Cascade algorithm by Yusheng Hu, completely avoiding store-forwarding stalls during hot loops. To scale this engine across multi-core processors, we extend the framework into a highly concurrent environment via OpenMP using a localized mathematical state decoder and macro-period loop scheduling. The parallel performance is shown to scale strictly and linearly with the number of active physical processor cores ($\text{Speedup}(M) \approx M$) due to a lock-free thread-local accumulation pipeline that eliminates false sharing. On a 6-core processor, the multi-threaded engine delivers a $5.25\times$ throughput gain for $n=14$ ($3.25$ billion CPU cycles) and processes the massive $n=15$ space in just $16.0$ seconds, yielding a $5.25\times$ speedup over the sequential vector baseline, while Hyper-Threading virtual cores yield zero additional throughput due to physical SIMD port saturation.

补充信息

↑