arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BRiDCT:使用SIMD的快速二维DCT:SIMD组织、寄存器分块与数值验证

BRiDCT: Fast Two-Dimensional DCTs Using SIMD: SIMD Organization, Register Blocking, and Numerical Verification

Antoine Moevus, Max Mignotte

arXiv 2609.28519首次发表:更新:

发表机构

Université de Montréal(蒙特利尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BRiDCT通过SIMD组织和寄存器分块实现快速二维DCT,在多种尺寸下超越主流库,并经过严格数值验证。

AI 中文摘要

重复的二维离散余弦变换(DCT)需要能够在多种数组尺寸下保持快速且不改变数学变换的实现。较少的算术运算并不能保证更快的DCT:执行时间还取决于计算如何使用CPU的单指令多数据(SIMD)单元、寄存器和内存层次结构。我们用SIMD实现了已有的DCT算法,并比较了在float32数组从8×8到256×256上的前向、逆和归一化往返变换的执行策略。由此产生的库BRiDCT结合了Shao–Johnson分解与SIMD通道组织、带状中间存储、填充和寄存器分块,使用C11/ARM NEON实现。在一台Apple M3 Max上,最终的原生库在8×8到256×256的所有可用情况下,执行时间低于测试的通用路径——Apple的Accelerate框架(vDSP)、FFTW和Ooura。它还优于测试的libjxl和浮点libjpeg-turbo路径。编译后的NumPy接口在每次会话的所有130个测试合同中优于OpenCV,包括高达1024×1024的方形和两个矩形,且包含了调用和分配成本。补充比较评估了SciPy、DUCC和缓存/计划的pyFFTW。测得的增益来自SIMD执行和中间存储的组织,而非引入新的算术分解。一个可复现的474输入验证语料库探测了抵消、相邻浮点数、动态范围和极端振幅:最终配置的所有1260个核心检查均通过,而边界诊断揭示了溢出和弱分量限制。较大的原生变换显示出与尺寸相关的限制。

英文摘要

Repeated two-dimensional discrete cosine transforms (DCTs) require implementations that remain fast across several array sizes while preserving the mathematical transform. Fewer arithmetic operations do not guarantee a faster DCT: execution time also depends on how the computation uses the CPU's single-instruction, multiple-data (SIMD) units, registers and memory hierarchy. We implement established DCT algorithms in SIMD and compare execution strategies for forward, inverse, and normalized round-trip transforms on float32 arrays from $8\times8$ to $256\times256$. The resulting library, BRiDCT, combines the Shao--Johnson factorization with SIMD lane organization, banded intermediate storage, padding, and register blocking in C11/ARM NEON. On one Apple M3 Max, the final native library has lower execution times than the tested general-purpose routes---Apple's Accelerate framework (vDSP), FFTW and Ooura---across their available cases from $8\times8$ to $256\times256$. It also outperforms the tested libjxl and floating-point libjpeg-turbo routes. The compiled NumPy interface outperforms OpenCV in all 130 tested contracts per session, including squares up to $1024\times1024$ and two rectangles, with call and allocation costs included. Complementary comparisons evaluate SciPy, DUCC and cached/planned pyFFTW. The measured gains come from organizing SIMD execution and intermediate storage, rather than introducing a new arithmetic factorization. A reproducible 474-input verification corpus probes cancellation, neighboring floats, dynamic range, and extreme amplitudes: all 1 260 core checks of the final configuration pass, while boundary diagnostics reveal overflow and weak-component limitations. Larger native transforms show size-dependent limits. Code, versioned releases, dataset and recorded measurements: https://github.com/antmoev/bridct.

CommentsTechnical report, 13 pages, 5 figures; supplementary methods and results, 7 pages, 1 figure. Code, verification data and recorded measurements included as ancillary archives

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑