arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

向量处理器上用于Transformer推理的屋顶线稀疏张量收缩

At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference

Bowen Wang, Chi Zhang, Diyou Shen, Renzo Andri, Navaneeth Kunhi Purayil, Luca Benini

arXiv 2607.25504首次发表:更新:

发表机构

ETH Future Computing Laboratory (EFCL); Huawei ZRC(ETH未来计算实验室; 华为ZRC)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对Transformer推理中稀疏张量收缩问题,提出Ventaglio运行时可配置稀疏执行单元,结合RVV ISA扩展,通过索引收集等支持提升性能,在多集群系统中加速模型推理,相比密集基线有显著提速。

AI 中文摘要

细粒度权重剪枝和激活稀疏化已成为降低Transformer模型推理计算和内存成本的有效方法。在中等稀疏度情况下,古斯塔夫森数据流通过元数据驱动的索引累加为利用向量处理器上的激活和权重稀疏性提供了自然执行模型。然而,现有RVV架构缺乏对此模式的原生支持。我们提出Ventaglio,一个运行时可配置的稀疏执行单元,通过索引收集-累加-散射支持使稀疏张量收缩达到屋顶线性能。集成到开源向量处理集群并在12nm FinFET中实现后,Ventaglio将稀疏张量收缩内核加速6.9至7.4倍,面积开销仅3.1%。我们构建了Ventaglio扩展的性能精确指令级模型并用于多集群系统性能分析。在具有40%-60%双稀疏性的DuoGPT剪枝LLaMA-3-8B模型中,Ventaglio在预填充和自回归解码期间分别比密集基线加速2.40至5.25倍和2.06至3.16倍。

英文摘要

Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference for Transformer models. In the moderate-sparsity regime, Gustavson's dataflow provides a natural execution model for exploiting both activation and weight sparsity on vector processors through metadata-driven indexed accumulation. However, existing RVV architectures lack native support for this pattern, forcing kernels to rely on software index decoding and L1-backed indexed memory operations that keep sparse tensor contractions far below their roofline performance bound. We present Ventaglio, a runtime-configurable sparse execution unit coupled with RVV ISA extensions that drives sparse tensor contractions toward their roofline through indexed gather-accumulate-scatter support. Integrated into an open-source vector processing cluster and implemented in 12nm FinFET, Ventaglio accelerates sparse tensor contraction kernels by $6.9\text{--}7.4\times$ over optimized RVV baselines, with only $3.1\%$ area overhead for a cluster of tightly-L1 coupled vector processing elements. We build a performance-accurate instruction-level model of the Ventaglio extension, calibrate it against RTL implementation, and leverage it for scale-out performance analysis on a large $4\times4$ multi-cluster system. Using a DuoGPT-pruned LLaMA-3-8B model with practical $40\text{--}60\%$ dual sparsity, Ventaglio achieves $2.40\text{--}5.25\times$ and $2.06\text{--}3.16\times$ speedup over dense baselines during prefill and autoregressive decoding, respectively.

Comments5 pages, 4 figures, 34th IFIP/IEEE International Conference on Very Large Scale Integration SoC (VLSI-SoC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑