arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SVRF:面向长向量架构的高效寄存器存储

SVRF: Efficient Register Storage for Long-Vector Architectures

Francesco Minervini, Lorenzo Deltetto, Osman Unsal, Adrian Cristal

arXiv 2610.07078首次发表:更新:

发表机构

Barcelona Supercomputing Center - BSC-CNS(巴塞罗那超级计算中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长向量架构中向量寄存器文件面积与能耗开销大的问题,提出SVRF微架构增强,将标量语义向量值存入专用标量寄存器文件,减少VRF需求,实现最高15.8%面积、12%功耗降低及1.12倍加速。

AI 中文摘要

向量处理器利用数据级并行提供高计算吞吐量,但非常长的向量会使向量寄存器文件(VRF)成为面积、能耗和寄存器压力的重要来源。当向量指令操作具有标量语义的值时,这种成本会加剧:尽管这些值在架构上使用向量寄存器表示,但它们不需要全宽存储。本文提出标量化向量寄存器文件(SVRF),一种微架构增强方案,将标量语义的向量值重定向到专用标量寄存器文件,避免不必要的全宽物理向量寄存器分配和访问。SVRF与向量寄存器重命名集成,同时保持对支持的向量ISA的兼容性。我们在基于寄存器传输级(RTL)的、支持非常长向量的RISC-V向量扩展(RVV)兼容向量处理单元中实现并评估SVRF。在代表性的HPC和AI/ML工作负载中,SVRF在不降低性能的情况下减少了向量寄存器压力和VRF活动。通过利用由此产生的VRF需求减少,所提出的设计在评估配置中实现了高达15.8%的VRF面积减少和高达6%的总面积减少,同时总功耗降低高达12%。某些工作负载还实现了高达1.12倍的加速。

英文摘要

Vector processors exploit data-level parallelism to provide high computational throughput, but very long vectors can make the Vector Register File (VRF) a significant source of area, energy, and register pressure. This cost is exacerbated when vector instructions operate on values with scalar semantics: although such values are architecturally represented using vector registers, they do not require full-width storage. This paper proposes the Scalarized Vector Register File (SVRF), a microarchitectural enhancement that redirects scalar-semantics vector values to a dedicated scalar register file, avoiding unnecessary allocation and access of full-width physical vector registers. The SVRF integrates with vector register renaming while preserving compatibility with the supported vector ISA. We implement and evaluate SVRF in an Register Transfer Level (RTL)-based RISC-V Vector Extension (RVV)-compliant vector processing unit supporting very long vectors. Across representative HPC and AI/ML workloads, the SVRF reduces vector-register pressure and VRF activity without degrading performance. By exploiting the resulting reduction in VRF demand, the proposed design achieves up to 15.8% VRF area reduction and up to 6% total area reduction, while reducing total power by up to 12% in evaluated configurations. Some workloads also achieve up to 1.12X speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑