arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02988cs.AReess.SP

CAMTA:用于非线性函数近似的可重构多区域激活单元

CAMTA: A Reconfigurable Multi-Region Activation Unit for Nonlinear Function Approximation

Carlos Soto-Porras, Jose Fonseca-Cruz, Pablo Ramirez-Morera, Erick Obregon-Fonseca, Luis G. Leon-Vega, Jorge Castro-Godinez

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出CAMTA可重构多区域激活单元,用于FPGA和ASIC加速器的非线性函数近似,经实验验证其误差低、资源占用合理,可复用多个非线性函数。

中文摘要 AI 辅助

非线性激活函数广泛应用于机器学习工作负载,但它们的直接硬件实现通常成本高昂、功能特定,或难以在不同模型间复用。本研究提出CAMTA,一种用于FPGA和ASIC加速器中非线性函数近似的16位可重构多区域激活单元。CAMTA在共享的基于霍纳(Horner)的数据路径上,结合了独立的区域阈值、各区域的多项式次数、系数集和执行模式。与传统的多项式或分段近似单元(主要重新配置系数或分段选择)不同,CAMTA还通过霍纳(HORNER)、常数(CONST)、零(ZERO)和恒等(IDENTITY)模式重新配置每个区域的计算行为,使同一硬件无需重新综合即可支持具有不同对称性和尾部行为的函数。在AMD Alveo平台上的FPGA验证显示,CAMTA辅助的Softmax的均方根误差(RMSE)低至3.60×10⁻⁶,比本研究中考虑的基于CORDIC的Softmax基线性能提升近一个数量级。FPGA高层次综合(HLS)综合报告显示,其使用3个数字信号处理器(DSP)、802个触发器(FF)、1756个查找表(LUT),数据路径延迟为11个周期。在台积电(TSMC)65nm工艺、250MHz下的ASIC综合报告显示,总单元面积为6632.40μm²,总功耗为1.3634mW。与同节点下功能特定的PLAC实现相比,CAMTA带来2.20倍的面积和1.75倍的功耗开销,但换来了运行时可配置性以及在多个非线性函数间的复用能力。

英文摘要

Nonlinear activation functions are widely used in machine learning workloads, but their direct hardware implementation is often costly, function-specific, or difficult to reuse across different models. This work introduces CAMTA, a 16-bit reconfigurable multi-region activation unit for nonlinear function approximation in FPGA and ASIC accelerators. CAMTA combines independent region thresholds, per-region polynomial degrees, coefficient sets, and execution modes over a shared Horner-based datapath. Unlike conventional polynomial or piecewise approximation units that mainly reconfigure coefficients or segment selection, CAMTA also reconfigures the computational behavior of each region through HORNER, CONST, ZERO, and IDENTITY modes, enabling the same hardware to support functions with different symmetry and tail behavior without resynthesis. FPGA validation on an AMD Alveo platform shows RMSE as low as \(3.60\times10^{-6}\) for CAMTA-assisted Softmax, outperforming the CORDIC-based Softmax baseline considered in this work by nearly one order of magnitude. FPGA HLS synthesis reports 3 DSPs, 802 FFs, 1756 LUTs, and an 11-cycle datapath latency. ASIC synthesis in TSMC 65~nm at 250~MHz reports \(6632.40~μ\mathrm{m}^2\) total cell area and \(1.3634~\mathrm{mW}\) total power. Compared with a same-node, function-specific PLAC implementation, CAMTA incurs \(2.20\times\) area and \(1.75\times\) power overhead, in exchange for runtime configurability and reuse across multiple nonlinear functions.

补充信息

↑