AI 中文总结
该研究针对FPGA上LLM非线性函数实现成本高的问题,提出基于部分重配置的PWL插值框架,通过动态加载模块实现面积节省,验证了分段策略的权衡效果。
AI 中文摘要
指数函数、Sigmoid等非线性函数是AI与大语言模型(LLM)加速的关键,但在现场可编程门阵列(FPGA)上高效实现仍成本高昂。本文提出一种基于部分重配置的分段线性(PWL)插值框架,以降低硬件成本同时保留灵活性。该架构将设计划分为用于通信与控制的静态区域,以及可动态加载不同插值模块的可重配置区域。针对指数函数与Sigmoid函数,采用FP16与FP32算术运算,评估了均匀与非均匀分段策略。结果显示,非均匀分段可提升高曲率区域的精度,而均匀分段的硬件开销更低。在系统层面,与包含上述两种算子的静态设计相比,该可重配置实现取得了显著的面积节省,查找表(LUT)减少最多达43%,触发器、块RAM(BRAM)及数字信号处理(DSP)单元减少最多达50%,且重配置延迟可预测。这些结果表明,部分重配置是探索FPGA上LLM工作负载非线性函数加速的面积-延迟权衡的可行方法。
英文摘要
Non-linear functions such as exponential and sigmoid are essential in AI and LLM acceleration, although implementing them efficiently on FPGAs is still costly. This paper proposes a PWL interpolation framework based on partial reconfiguration to reduce hardware cost while preserving flexibility. The architecture separates the design into a static region for communication and control, and a reconfigurable region where different interpolation modules can be dynamically loaded. Uniform and non-uniform segmentation strategies are evaluated for exponential and sigmoid functions using FP16 and FP32 arithmetic. Results show that non-uniform segmentation can improve accuracy in high-curvature regions, while uniform segmentation offers lower hardware overhead. At the system level, the reconfigurable implementation achieved significant area savings, reaching up to 43\% less LUTs, 50\% less flip-flops, BRAMs and DSPs cells, compared against a static design containing both operators; all this with predictable reconfiguration latency. These results show that partial reconfiguration is a practical approach for exploring area-latency trade-offs in FPGA-based acceleration of non-linear functions for LLM workloads.
CommentsSubmitted to ICECS 2026