arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多核CPU和GPU上的并行级联递归滤波

Parallel Cascaded Recursive Filtering on Multi-Core CPUs and GPUs

Haotian Zhai, Bernd-Peter Paris

arXiv 2607.23763首次发表:更新:

AI 中文总结

研究将级联二阶递归滤波框架扩展到多核CPU和GPU时的问题,通过叠加和分治循环约简解决依赖性,给出实时流和批处理两种实现方式,在性能上有显著提升,超过已发表基线且数值稳定性好。

AI 中文摘要

本文的配套论文将级联二阶(双二阶)递归滤波重新表述为块三对角线性系统,并开发了两种并行求解算法,PH分解和循环约简,在单个SIMD核心上达到每秒600兆样本以上。本文将该框架扩展到多核CPU和GPU,出现了新障碍:每个信号块组的终端输出是下一组的初始条件,天真分布的组会串行化。通过叠加(每个组的输出分为可立即计算的零状态响应和状态到达时应用的齐次校正)和一种分治形式的循环约简(在回代之前暴露两个终端块,因为异步状态传播需要)解决依赖性。两种实现将两种主要部署场景与对依赖性的相反处理配对。对于实时流,用TBB流程图实现的波前流水线跨级联部分并行化,保持先进先出顺序,在六个性能核心上实现3.95倍的扩展,对于16阶滤波器约为每秒2.4吉样本。对于批处理,单核GPU实现通过寄存器将每个组带入整个级联,并使用解耦回查协议跨组并行化;基于通信的成本模型(包括内存上限、屏障代价和延迟隐藏下限)将调优减少到两个参数,并预测两代GPU的测量行为。最好的内核在RTX~3060上对于单个二阶部分达到每秒38.2吉样本,是内存带宽上限的85%,在每个滤波器阶数上超过最强的已发表并行递归基线,并且在16阶时仍然数值有效,而直接形式基线失败。

英文摘要

The companion of this paper reformulated cascaded second-order (biquad) recursive filtering as a block-tridiagonal linear system and developed two parallel solution algorithms, PH factorization and cyclic reduction, reaching over 600 Megasamples per second on a single SIMD core. This paper scales that framework to multi-core CPUs and GPUs, where a new obstacle appears: the terminal outputs of each signal block group are the initial conditions of the next, so naively distributed groups serialize. The dependency is resolved by superposition -- each group's output splits into a zero-state response, computable immediately, and a homogeneous correction applied when the state arrives -- and by a divide-and-conquer form of cyclic reduction that exposes both terminal blocks before back substitution, as asynchronous state propagation requires. Two implementations pair the two dominant deployment scenarios with opposite treatments of the dependency. For real-time streaming, a wavefront pipeline realized with TBB flow graphs parallelizes across cascade sections, preserves first-in-first-out order, and achieves 3.95x scaling on six performance cores, about 2.4 Gigasamples per second for a 16th-order filter. For batched processing, a single-kernel GPU implementation carries each group through the entire cascade in registers and parallelizes across groups with a decoupled look-back protocol; a communication-based cost model, comprising a memory roof, a barrier price, and a latency-hiding floor, reduces tuning to two parameters and predicts the measured behavior across two GPU generations. The best kernels reach 38.2 Gigasamples per second for a single second-order section on an RTX~3060, 85% of the memory-bandwidth roof, exceed the strongest published parallel recurrence baseline at every filter order, and remain numerically valid at order 16, where the direct-form baseline fails.

Comments14 pages, 13 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑