发表机构
Nara Institute of Science and Technology(奈良先端科学技术大学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文面向存储多项式数字预失真核,将其映射到IMAX CGLA架构,实现低延迟与高能效,相比RTX 4090能耗降低169.1倍,同时保持线性化性能。
AI 中文摘要
存储多项式数字预失真(DPD)在滑动输入历史上评估一个小的固定系数集,因此其归约步骤是一个具有局部复用的复数乘累加(complex-MAC)工作负载。我们将该DPD归约核映射到内存加速器扩展(IMAX)上,IMAX是一种可编程的CPU接地线性阵列(CGLA),由一维处理单元/本地存储器流水线组成。对于(P,M)=(5,5)的奇阶存储多项式实例,该映射将120字节系数集保存在本地存储器中,在1024样本瓦片上推进五抽头历史,并将15个阶延迟项实现为33级流式复数乘累加归约。评估测量了核延迟和建模能量。所有测量路径使用相同的单精度复数工作负载,即32个序列,每个序列2048个复数样本,分别在IMAX FPGA原型、RTX 4090系统上的CUDA实现以及Jetson AGX Orin上的ARM-NEON实现上进行。使用此1024样本瓦片配置,IMAX FPGA原型报告端到端延迟为20.201毫秒,仅核延迟为1.948毫秒。使用先前报告的28纳米IMAX频率和功率模型,预计IMAX配置给出3.14毫秒的端到端延迟和0.34毫秒的仅核延迟。RTX 4090基线具有最低的端到端延迟,为0.484毫秒。在基于模型的平台功率核算和所述功率假设下,预计IMAX配置每批次的建模端到端能量比RTX 4090基线小169.1倍。该值使用平台功率假设,而非依赖于工作负载的运行时功率或直接硅片功率测量。受控合成PA模型验证检查了相同的15项形式将测试集NMSE提高26.1分贝,ACLR提高26.0分贝。这些结果表征了在评估的瓦片配置和功率模型下,IMAX上映射的存储多项式DPD归约。
英文摘要
Memory-polynomial digital predistortion (DPD) evaluates a small fixed coefficient set over a sliding input history, so its reduction step is a complex-MAC workload with local reuse. We map this DPD reduction kernel onto In-Memory Accelerator eXtension (IMAX), a programmable CPU-Grounded Linear Array (CGLA) composed of a one-dimensional processing-element/local-memory pipeline. For a (P,M)=(5,5) odd-order memory-polynomial instance, the mapping keeps the 120 B coefficient set in local memory, advances the five-tap history over 1024-sample tiles, and realizes the 15 order-delay terms as a 33-stage streaming complex-MAC reduction. The evaluation measures kernel latency and modeled energy. All measured paths use the same single-precision complex workload of 32 sequences, each with 2048 complex samples, across an IMAX FPGA prototype, a CUDA implementation on an RTX 4090 system, and an ARM-NEON implementation on Jetson AGX Orin. With this 1024-sample tile configuration, the IMAX FPGA prototype reports 20.201 ms end-to-end latency and 1.948 ms kernel-only latency. Using the previously reported 28 nm IMAX frequency and power model, the projected IMAX configuration gives 3.14 ms end-to-end latency and 0.34 ms kernel-only latency. The RTX 4090 baseline has the lowest end-to-end latency at 0.484 ms. Under model-based platform power accounting and the stated power assumptions, the projected IMAX configuration gives 169.1 times smaller modeled end-to-end energy per batch than the RTX 4090 baseline. This value uses platform power assumptions rather than workload-dependent runtime power or a direct silicon power measurement. A controlled synthetic PA-model validation checks that the same 15-term form improves test-set NMSE by 26.1 dB and ACLR by 26.0 dB. These results characterize the mapped memory-polynomial DPD reduction on IMAX for the evaluated tile configuration and power model.
CommentsAccepted for presentation at the 2026 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS 2026)