一种用于字典序排列的分治引擎:通过软硬件协同的CPU指令加速状态演化
A Divide-and-Conquer Engine for Lexicographical Permutations: Accelerating State Evolution via Hybrid Software-Hardware CPU Instructions
浏览论文内容
中文总结 AI 辅助
该研究提出LexCHA软硬件协同架构,利用SIMD指令和分形同构性实现无分支排列执行,大幅提升字典序排列的吞吐量,优于现有std::next_permutation实现。
中文摘要 AI 辅助
以std::next_permutation为代表的传统字典序排列算法,根本上受限于密集控制逻辑和频繁的分支预测错误。本文提出LexCHA(字典序协同硬件加速),一种新颖的软硬件协同架构,通过原生SIMD指令加速排列。利用排列固有的分形同构性,LexCHA将全局状态演化解耦为宏观软件排秩阶段和微观硬件块构建阶段。通过利用预计算的确定性掩码流,LexCHA用流式向量洗牌替代传统条件分支,实现无分支执行流。评估显示,LexCHA比现代处理器的标量实现吞吐量显著更高,大幅优于现有最先进的std::next_permutation实现。
英文摘要
Traditional lexicographical permutation algorithms, epitomized by \texttt{std::next\_permutation}, are fundamentally bottlenecked by dense control logic and frequent branch mispredictions. This paper introduces LexCHA(Lexicographical Co-designed Hardware Acceleration), a novel hardware-software co-designed architecture that accelerates permutation via native SIMD instructions. Exploiting the inherent fractal isomorphism of permutations, LexCHA decouples global state evolution into a macro software unranking phase and a micro hardware block-construction phase. By utilizing pre-computed deterministic mask streams, LexCHA replaces traditional conditional branching with streaming vector shuffles, achieving a branch-free execution flow. Evaluations demonstrate that LexCHA demonstrates significantly higher throughput than scalar implementations of modern processors, outperforming existing state-of-the-art \texttt{std::next\_permutation} implementations significantly.