发表机构
College of Computer Science and Artificial Intelligence, Fudan University; School of Mathematical Sciences, School of Data Science, Fudan University; Shanghai Key Lab. of Intelligent Information Processing(复旦大学计算机科学与人工智能学院; 复旦大学数学科学学院、数据科学学部; 上海市智能信息处理重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PACO提出一种完全缓存无关的并行FFT框架,通过延迟递归布局变换并融合为一次全局数字反转排列,实现单次全局重分布下的最优缓存与通信复杂度。
AI 中文摘要
并行机器上的快速傅里叶变换涉及两种形式的数据移动:通过内存层次结构的处理器本地传输和处理器之间的全局重分布。四步FFT组织通过一次全局转置类交换来切换活动变换维度,但这本身并不能产生缓存高效的本地计算。缓存无关FFT通过递归布局变换实现渐近最优的本地内存流量,其直接物化可能需要额外的数据重排步骤和全局交换。我们提出PACO,一个完全缓存无关的并行FFT框架,调和了这些目标。PACO执行LocalFFT -> OneGlobalPermutation -> LocalFFT。其本地阶段递归地划分变换维度和批处理维度,无需知道缓存参数。PACO不物化由该递归引起的转置类布局,而是延迟它们,表明它们组合成一个基b数字反转排列,该排列与改变本地变换维度所需的重新分布融合在一起。由此产生的中间阶段是一个完美平衡的并行缓存无关数字反转排列,其中每个源-目标处理器对恰好交换$N/p^2$个元素。在混合理想缓存/BSP模型中的精确基b切片分解下,PACO精确计算N点DFT,每处理器最大工作量$\theta((N/p) \n log N)$,每处理器最大缓存复杂度$\theta((N/(pB))(1 + \n log_M N))$,恰好使用一轮全局重分布。PACO在因子交换的目标切片所有权下返回规范逻辑DFT系数。在所声明的无复制所有权模型下,单次重分布是必要的,而其通信量对于规定的融合排列和源-目标切片分布是最优的。
英文摘要
Fast Fourier transforms on parallel machines incur two forms of data movement: processor-local transfers through the memory hierarchy and global redistribution between processors. Four-step FFT organizations switch the active transform dimension with one global transpose-like exchange, but this alone does not yield cache-efficient local computation. Cache-oblivious FFTs achieve asymptotically optimal local memory traffic through recursive layout transformations, whose direct materialization can require extra data-rearrangement passes and global exchanges. We present PACO, a fully cache-oblivious parallel FFT framework that reconciles these objectives. PACO executes LocalFFT -> OneGlobalPermutation -> LocalFFT. Its local stages recursively partition both transform and batch dimensions without knowledge of the cache parameters. Rather than materializing the transpose-like layouts induced by this recursion, PACO defers them, showing that they compose into a base-b digit-reversal permutation that is fused with the redistribution already required to change the local transform dimension. The resulting middle stage is a perfectly balanced parallel cache-oblivious digit-reversal permutation in which each source-destination processor pair exchanges exactly $N/p^2$ elements. Under an exact base-b slab decomposition in a hybrid ideal-cache/BSP model, PACO computes an N-point DFT exactly with maximum per-processor work $Θ((N/p) \log N)$ and maximum per-processor cache complexity $Θ((N/(pB))(1 + \log_M N))$, using exactly one global redistribution round. PACO returns canonical logical DFT coefficients under a factor-swapped target slab ownership. The single redistribution is necessary under the stated no-replication ownership model, while its communication volume is optimal for the prescribed fused permutation and source-target slab distributions.
Comments11 pages, 1 figure, with about 10 appendices as supplementary materials