arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向存储多项式数字预失真核的能量导向CGLA映射

Energy-Oriented CGLA Mapping of a Memory-Polynomial Digital Predistortion Kernel

Takuto Ando, Yasuhiko Nakashima

arXiv 2609.27438首次发表:更新:

发表机构

Nara Institute of Science and Technology(奈良先端科学技术大学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文面向存储多项式数字预失真核,将其映射到IMAX CGLA架构,实现低延迟与高能效,相比RTX 4090能耗降低169.1倍,同时保持线性化性能。

AI 中文摘要

存储多项式数字预失真(DPD)在滑动输入历史上评估一个小的固定系数集,因此其归约步骤是一个具有局部复用的复数乘累加(complex-MAC)工作负载。我们将该DPD归约核映射到内存加速器扩展(IMAX)上,IMAX是一种可编程的CPU接地线性阵列(CGLA),由一维处理单元/本地存储器流水线组成。对于(P,M)=(5,5)的奇阶存储多项式实例,该映射将120字节系数集保存在本地存储器中,在1024样本瓦片上推进五抽头历史,并将15个阶延迟项实现为33级流式复数乘累加归约。评估测量了核延迟和建模能量。所有测量路径使用相同的单精度复数工作负载,即32个序列,每个序列2048个复数样本,分别在IMAX FPGA原型、RTX 4090系统上的CUDA实现以及Jetson AGX Orin上的ARM-NEON实现上进行。使用此1024样本瓦片配置,IMAX FPGA原型报告端到端延迟为20.201毫秒,仅核延迟为1.948毫秒。使用先前报告的28纳米IMAX频率和功率模型,预计IMAX配置给出3.14毫秒的端到端延迟和0.34毫秒的仅核延迟。RTX 4090基线具有最低的端到端延迟,为0.484毫秒。在基于模型的平台功率核算和所述功率假设下,预计IMAX配置每批次的建模端到端能量比RTX 4090基线小169.1倍。该值使用平台功率假设,而非依赖于工作负载的运行时功率或直接硅片功率测量。受控合成PA模型验证检查了相同的15项形式将测试集NMSE提高26.1分贝,ACLR提高26.0分贝。这些结果表征了在评估的瓦片配置和功率模型下,IMAX上映射的存储多项式DPD归约。

英文摘要

Memory-polynomial digital predistortion (DPD) evaluates a small fixed coefficient set over a sliding input history, so its reduction step is a complex-MAC workload with local reuse. We map this DPD reduction kernel onto In-Memory Accelerator eXtension (IMAX), a programmable CPU-Grounded Linear Array (CGLA) composed of a one-dimensional processing-element/local-memory pipeline. For a (P,M)=(5,5) odd-order memory-polynomial instance, the mapping keeps the 120 B coefficient set in local memory, advances the five-tap history over 1024-sample tiles, and realizes the 15 order-delay terms as a 33-stage streaming complex-MAC reduction. The evaluation measures kernel latency and modeled energy. All measured paths use the same single-precision complex workload of 32 sequences, each with 2048 complex samples, across an IMAX FPGA prototype, a CUDA implementation on an RTX 4090 system, and an ARM-NEON implementation on Jetson AGX Orin. With this 1024-sample tile configuration, the IMAX FPGA prototype reports 20.201 ms end-to-end latency and 1.948 ms kernel-only latency. Using the previously reported 28 nm IMAX frequency and power model, the projected IMAX configuration gives 3.14 ms end-to-end latency and 0.34 ms kernel-only latency. The RTX 4090 baseline has the lowest end-to-end latency at 0.484 ms. Under model-based platform power accounting and the stated power assumptions, the projected IMAX configuration gives 169.1 times smaller modeled end-to-end energy per batch than the RTX 4090 baseline. This value uses platform power assumptions rather than workload-dependent runtime power or a direct silicon power measurement. A controlled synthetic PA-model validation checks that the same 15-term form improves test-set NMSE by 26.1 dB and ACLR by 26.0 dB. These results characterize the mapped memory-polynomial DPD reduction on IMAX for the evaluated tile configuration and power model.

CommentsAccepted for presentation at the 2026 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑