发表机构
University College Dublin; Trinity College Dublin; Virginia Tech; Queen’s University Belfast(都柏林大学学院; 都柏林圣三一学院; 弗吉尼亚理工大学; 贝尔法斯特女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本框架通过指令集扩展和循环展开策略,在RISC-V边缘设备上加速紧凑型Transformer模型,实现最高2.19倍推理加速,同时保持较低的硬件开销。
AI 中文摘要
本工作提出了一种在资源受限的物联网设备上加速基于Transformer的语言模型(LMs)的框架。该框架针对紧凑型语言模型:BERT-Tiny (B-Ty)、MobileBERT (M-Bt)、MiniLM (M-Lm)、Electra (E-Lt) 和 DeBERTa (D-Bt)——这些模型因其架构多样性和在边缘推理场景中的使用而被选中。所提出的流程针对这些模型中不明显的计算模式,衍生出轻量级指令集扩展。此外,引入了一条自定义指令来加速批量矩阵乘法的地址生成阶段,实现了15.39%–21.74%的性能提升,同时ASIC开销仅为面积6.79%和功耗2.33%。为了在不增加额外处理器核心硬件成本的情况下进一步提升性能,采用了一种可选的编译器引导的循环展开策略,以增加代码大小为代价换取整体执行时间的减少。在Synopsys trv32p3f RISC-V核心上的评估显示,推理加速最高可达2.19倍;在AMD Zynq UltraScale+ ZCU102上的FPGA实现显示,在75 MHz下面积开销为32.94%,而使用TSMC 28 nm工艺的ASIC实现,在250 MHz下面积开销为21.08%。
英文摘要
This work presents a framework for accelerating transformer-based language models (LMs) on resource-constrained IoT devices. The framework targets compact LMs: BERT-Tiny (B-Ty), MobileBERT (M-Bt), MiniLM (M-Lm), Electra (E-Lt) and DeBERTa (D-Bt) -- selected for their architectural diversity and use in edge inference scenarios. The proposed flow derives lightweight instruction set extensions tailored to the non-obvious computational patterns of these models. In addition, a custom instruction is introduced to accelerate the address generation stage of batch matrix multiplication, achieving a performance improvement of 15.39--21.74% with modest ASIC overheads of 6.79% in area and 2.33% in power. To further enhance performance without incurring additional processor core hardware cost, an optional compiler-directed loop unrolling strategy is employed, trading increased code size for overall reduced execution time. Evaluation on the Synopsys trv32p3f RISC-V core demonstrates inference speedups of up to 2.19x, and FPGA implementation on the AMD Zynq UltraScale+ ZCU102 shows a 32.94% area overhead at 75 MHz, whereas the ASIC implementation using the TSMC 28 nm library incurs a 21.08% area overhead while operating at 250 MHz.
CommentsAccepted for publication at the 2026 IEEE 33rd International Conference on Electronics, Circuits and Systems (ICECS)