AI 中文总结
研究光正交频分复用收发器中FFT核数字格式,提出跨层协同设计方法,联合优化数字表示等。以128 Gbit/s收发器中256点FFT引擎为例,比较定点和低精度浮点格式,新格式保持误码率性能,降低功耗和面积。
AI 中文摘要
许多先进的数字信号处理(DSP)实现采用定点算法,因其硬件复杂度低且吞吐量高。相比之下,机器学习加速器采用低精度浮点格式能显著提高能效和峰值吞吐量。这促使重新评估经典DSP工作负载的数字表示。核心问题是低精度浮点格式能否在功耗、性能和面积上与定点实现竞争,同时在动态范围和数值鲁棒性上具有优势。本文提出一种新的跨层协同设计方法,联合优化数字表示、算术单元和应用级性能。以FFT为例,在光正交频分复用收发器中评估,FFT在其中占主导地位。比较定点和低精度浮点格式,展示了11位和12位自定义浮点格式在保持误码率性能的同时,降低了FFT核心功耗和面积。据我们所知,这是首次对光正交频分复用收发器中FFT核的自定义低精度浮点算法进行研究。
英文摘要
Many state-of-the-art DSP implementations use fixed-point arithmetic due to its reduced hardware complexity and high throughput compared to conventional floating-point arithmetic. In contrast, machine learning accelerators exhibit substantial gains from reduced-precision floating-point formats, enabling improvements in energy efficiency and peak throughput. These advances motivate a re-evaluation of numerical representations for classical DSP workloads. A central question is whether reduced-precision floating-point formats can achieve competitive power, performance, and area compared to fixed-point implementations, while providing advantages in dynamic range and numerical robustness. This paper presents a new cross-layer co-design methodology for DSP kernels that jointly optimizes numerical representations, arithmetic units, and application-level performance. As a case study, we focus on the FFT, a fundamental DSP block across many applications. The FFT is evaluated within optical OFDM transceivers, where it dominates power consumption and silicon area as FFT size and modulation order scale to support data rates beyond 100 Gbit/s. We compare fixed-point and reduced-precision floating-point formats using post-layout power and area results in a 12nm FinFET technology and demonstrate system-level performance in terms of BER versus Eb/N0. For a 256-point FFT engine in a 128 Gbit/s transceiver, we show that 11- and 12-bit custom floating-point formats preserve BER performance close to a 32-bit floating-point reference across multiple modulation orders, while reducing FFT core power by up to 19.8% and area by up to 12.0% compared to representative fixed-point designs. To the best of our knowledge, this is the first investigation of custom reduced-precision floating-point arithmetic for FFT cores in optical OFDM transceivers.
CommentsAccepted for presentation at the 2026 IEEE International Workshop on Signal Processing Systems (SiPS 2026). Proceedings to be included in IEEE Xplore