AI 中文总结
研究针对传统软件CRC实现的顺序依赖性问题,提出基于POSIX线程的多核架构通用软件并行CRC框架,支持多种CRC变体,采用特定组合机制,实验表明其显著提升大数据集性能,提供了可移植的正确性并行CRC加速框架。
AI 中文摘要
循环冗余校验(CRC)仍是通信、存储和嵌入式系统中最广泛使用的错误检测机制之一。然而,传统软件CRC实现存在固有的顺序依赖性,限制了现代多核处理器的有效利用。本文提出了一种基于POSIX线程的多核架构通用软件并行CRC框架。该框架在统一实现模型中支持多种CRC变体,包括CRC-8、CRC-16、CRC-32、CRC-64和CRC-128。为在并行执行期间保持正确性,框架采用基于GF(2)的CRC组合机制而非简单的异或聚合。组合阶段使用GF(2)上的多项式算术和基于矩阵的移位操作来制定,确保并行和串行CRC计算的等效性。使用多种工作负载大小和线程配置对该方法进行了评估。实验分析包括不同线程数下的执行时间、吞吐量、延迟、可扩展性行为和能量估计。结果表明,并行执行显著提高了大数据集的性能,在评估平台上实现了约3至4倍的加速,同时保持了精确的CRC正确性。还与代表性的CRC优化方法进行了比较讨论,以将所提出的框架置于更广泛的CRC优化领域中。总体而言,该方法为通用多核系统上保持正确性的并行CRC加速提供了一个可移植的通用软件框架。
英文摘要
Cyclic Redundancy Check (CRC) remains one of the most widely used error-detection mechanisms in communication, storage, and embedded systems. However, conventional software CRC implementations suffer from inherent sequential dependencies that limit efficient utilization of modern multi-core processors. This paper presents a generalized software-based parallel CRC framework for multi-core architectures using POSIX threads. The proposed framework supports multiple CRC variants, including CRC-8, CRC-16, CRC-32, CRC-64, and CRC-128, within a unified implementation model. To preserve correctness during parallel execution, the framework employs a GF(2)-based CRC combination mechanism rather than naive XOR aggregation. The combine stage is formulated using polynomial arithmetic and matrix-based shifting operations over GF(2), ensuring equivalence between parallel and serial CRC computation. The proposed method was evaluated using multiple workload sizes and thread configurations. Experimental analysis includes execution time, throughput, latency, scalability behavior, and energy estimation under varying thread counts. Results indicate that parallel execution significantly improves performance for large datasets, achieving approximately 3-4x speedup on the evaluated platform while preserving exact CRC correctness. Comparative discussion with representative CRC optimization approaches, including lookup-table methods, slicing-by-8, SIMD/vectorized CRC, and hardware-assisted CRC techniques, is also provided to position the proposed framework within the broader CRC optimization landscape. Overall, the proposed approach provides a portable and generalized software framework for correctness-preserving parallel CRC acceleration on general-purpose multi-core systems.