AI 中文总结
研究递归滤波器并行计算问题,通过将双二阶差分方程重述为块托普利兹线性系统,引入特定置换,开发部分LU分解和循环约简两种并行算法,减少依赖深度,消除冗余阶段,提升计算效率,新算法在实验中表现优异。
AI 中文摘要
递归(IIR)滤波器实现为级联二阶节(双二阶滤波器),兼具设计通用性和抗系数量化鲁棒性。但其固有的逐样本反馈依赖性阻碍并行计算。本文将双二阶差分方程重新表述为带状块托普利兹线性系统,引入步长为\(N\)的置换,将一组\(NL\)个样本映射为块三对角结构。在此框架下开发了两种并行算法:保留稀疏块结构的部分LU(PH)分解和首次应用于递归滤波的循环约简,将顺序依赖深度从\(\mathcal{O}(N)\)降至\(\mathcal{O}(\log_2 N)\)。对于\(K\)个双二阶滤波器的级联,连续节之间的中间置换恰好抵消,整个级联仅需一对置换/反置换,消除\(2(K - 1)\)个冗余阶段。推导了各算法阶段的精确块级操作计数,并在三种支持AVX2 SIMD指令的英特尔微架构上通过精确周期测量进行验证。16阶系统的实验结果表明,与标量滤波相比拟议的多块算法将每样本时钟周期减少多达\(10\)倍,两种算法在新架构上扩展性良好。在单个Meteor Lake核心上,循环约简实现约618 MS/s,吞吐量比此http URL提高\(8\)倍。
英文摘要
Recursive (IIR) filters realized as cascaded second-order sections (biquads) offer both design generality and robustness against coefficient quantization. However, their inherent sample-to-sample feedback dependency poses a fundamental obstacle to parallel computation. This paper reformulates the biquad difference equation as a banded block-Toeplitz linear system and introduces a stride-$N$ permutation that maps a group of $NL$ samples into a block-tridiagonal structure whose entries are scalar multiples of identity and shift matrices. Within this framework, two parallel algorithms are developed for the recursive solution: a partial LU (PH) factorization that preserves the sparse block structure and a cyclic reduction that is applied to recursive filtering, to the best of our knowledge, for the first time. It reduces the sequential dependency depth from $\mathcal{O}(N)$ to $\mathcal{O}(\log_2 N)$. For a cascade of $K$ biquads, the intermediate permutations between successive sections cancel exactly, so that only a single permutation/de-permutation pair is required for the entire cascade, eliminating $2(K{-}1)$ redundant stages. Exact block-level operation counts are derived for every algorithmic stage and validated against cycle-accurate measurements on three Intel micro-architectures supporting AVX2 SIMD instructions. Experimental results for a 16th-order system show that the proposed multi-block algorithms reduce clock cycles per sample by up to $10\times$ compared to scalar filtering, with both algorithms scaling favorably on newer architectures. On a single Meteor Lake core, cyclic reduction achieves approximately 618 MS/s -- an $8\times$ throughput improvement over scipy.signal.sosfilt.
Comments13 pages, 8 figures