COBALT:用于LSTM任务的列交换优化位串行加速器
COBALT: Column-swapping Optimized Bit-serial Accelerator for LSTM Tasks
AI总结:
针对边缘设备LSTM部署的计算资源瓶颈,提出COBALT位串行压缩LSTM加速器,通过列交换等技术实现高权重压缩率,在TIMIT等基准上保持准确率且FPGA效率优于现有方案。
AI中文摘要:
长短期记忆(LSTM)网络被广泛部署在边缘设备上以处理实时序列任务,但其计算需求对资源受限硬件的部署构成挑战。本研究提出COBALT,一种基于标准循环矩阵向量乘法(MVM)的矩阵(输入)-向量(权重)重构与偏移二进制编码构建的位串行压缩LSTM加速器。一种新颖的列交换方案在成对生成的部分积(PP)的位级运行,系统性地暴露输出行间的冗余以减少PP生成器和选择器的数量。该冗余通过轻量校正单元进一步利用,该校正单元直接从配对行推导某行的输出。此外,对于块循环MVM,重定位各子MVM的移位-累加与校正单元可进一步降低资源使用。该压缩网络在TIMIT和LibriSpeech-100h基准上保持准确率的同时,实现了高达93.6%的权重压缩率。在现场可编程门阵列上,COBALT相对于最先进的LSTM加速器实现了更优的整体效率。
英文摘要:
Long Short-Term Memory (LSTM) networks continue to be widely deployed for real-time sequential tasks on edge devices, yet their computational demands challenge deployment on resource-constrained hardware. This work introduces COBALT, a bit-serial compressed LSTM accelerator built on a matrix (input)-vector (weight) reformulation of the standard circulant matrix-vector multiplication (MVM) with offset-binary coding. A novel column-swapping scheme operates on partial products at the bit level when generated in pairs, systematically exposing redundancy across output rows to reduce the number of PP generators and selectors. This redundancy is further exploited using a lightweight correction unit that derives a row's output directly from its paired row. Additionally, for block-circulant MVMs, relocating the shift-accumulate and correction units of each sub-MVM further reduces resource usage. The compressed network achieves weight compression of up to 93.6% while maintaining accuracy on the TIMIT and LibriSpeech-100h benchmarks. On a field-programmable gate array, COBALT achieves superior overall efficiency relative to state-of-the-art LSTM accelerators.