arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19431cs.ARcs.AI

BRIM:工作负载平衡的双边位串行稀疏推理加速器

BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator

Varun Manjunath, Ruokai Yin, Donghyun Lee, Arkapravo Ghosh, Priyadarshini Panda

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对双边位串行稀疏推理加速器工作负载不平衡问题,提出BRIM。它采用循环平衡剪枝和成对时隙捐赠两种机制,在等面积约束下评估,能提高PE利用率,实现显著加速和能效提升。

中文摘要 AI 辅助

位串行加速器利用位级稀疏性来降低深度神经网络推理成本,但现有设计仅在一个操作数上利用稀疏性,限制了加速效果。将稀疏性利用扩展到两个操作数可同时减少部分积,但会引入工作负载不平衡这一关键瓶颈。我们展示了现有双边设计中处理单元(PE)利用率仅为56%-64%。我们提出了BRIM,一种软硬件协同设计的双边位串行稀疏加速器,它直接针对这一瓶颈。BRIM结合了两种集成机制:1)循环平衡剪枝(CBP),一种训练后权重优化方法,根据分析的激活统计信息重塑权重表示,以离线均衡并发处理对之间的预期工作负载;2)成对时隙捐赠,一种轻量级硬件机制,以可忽略的面积开销吸收剩余的运行时不平衡。在等面积约束下对卷积神经网络(CNNs)、视觉Transformer(ViTs)和语言模型(LLMs)进行评估,BRIM的PE利用率超过90%,比先前的双边设计加速高达2.37倍,能效提高高达1.63倍。

英文摘要

Bit-serial accelerators exploit bit-level sparsity to reduce DNN inference cost, but existing designs exploit sparsity on only one operand, bounding the speedup. Extending sparsity exploitation to both operands simultaneously yields compounding reductions in partial products but introduces a critical new bottleneck: workload imbalance. Because each concurrent weight - activation pair's execution cost depends on the product of two independently varying operand non-zero bit counts, pairs that must complete together finish at vastly different times, leaving faster computations idle. We show this limits PE utilization to 56 - 64% in existing dual-sided designs. We present BRIM, a hardware - software co-designed dual-sided bit-serial sparse accelerator that directly targets this bottleneck. BRIM combines two integrated mechanisms: 1) Cyclic-Balanced Pruning (CBP), a post-training weight optimization that reshapes weight representations based on profiled activation statistics to equalize expected workloads across concurrently processed pairs offline; and 2) Pairwise Slot Donation, a lightweight hardware mechanism that absorbs residual runtime imbalance with negligible area overhead. Evaluated across CNNs, ViTs, and LLMs under iso-area constraints, BRIM achieves over 90% PE utilization, up to 2.37x speedup, and up to 1.63x energy efficiency improvement over prior dual-sided designs.

发表机构

  • University of Southern California(南加州大学)
  • Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

↑